Lobhari Tech Journal & Insights

Engineering Insights & Stories

Practical knowledge, architecture deep-dives, and startup strategy from the engineering team at Lobhari Technologies.

Showing 34 articles
Sort:
GLM-5.3 & The Rise of Cyber-Capable AI: Automated Vulnerability Discovery & Defensive ArchitectureCybersecurity
CybersecurityZhipu AI

GLM-5.3 & The Rise of Cyber-Capable AI: Automated Vulnerability Discovery & Defensive Architecture

An architectural deep dive into Zhipu AI's GLM-5.3: a 743B open-weight foundation model scoring 84.5% on CyberGym, automating zero-day vulnerability discovery, and reshaping enterprise DevSecOps.

Aug 21, 20265 min read
Inside Miles v0.1: The Production-Grade RL Post-Training Architecture for Trillion-Parameter MoE ModelsReinforcement Learning
Reinforcement LearningLMSYS

Inside Miles v0.1: The Production-Grade RL Post-Training Architecture for Trillion-Parameter MoE Models

An architectural deep dive into Miles v0.1 by LMSYS and RadixArk: scaling asynchronous RL post-training for 700B+ MoE models using SGLang rollouts, Rollout Routing Replay (R3), and Megatron-LM.

Aug 21, 20266 min read
DeepSeek V4-Pro GA & V4-Flash: Inside the 1.6T MoE Architecture, 384K Output Limits & Agent EconomicsDeepSeek
DeepSeekAI Models

DeepSeek V4-Pro GA & V4-Flash: Inside the 1.6T MoE Architecture, 384K Output Limits & Agent Economics

An architectural deep dive into DeepSeek V4-Pro GA (1.6T MoE) and V4-Flash: 384K output token limits, 87.9% Terminal-Bench scores, and peak/off-peak agent pricing economics.

Aug 21, 20265 min read
Inside Thinking Machines Inkling: Mira Murati's 975B Open-Weight MoE Foundation ModelThinking Machines
Thinking MachinesInkling

Inside Thinking Machines Inkling: Mira Murati's 975B Open-Weight MoE Foundation Model

An architectural deep dive into Thinking Machines Lab's Inkling: a 975B sparse Mixture-of-Experts foundation model activating 41B parameters per query, engineered by Mira Murati's team for enterprise open-weight autonomy.

Aug 21, 20265 min read
Beyond Naive RAG: Why Cross-Encoder Reranking & Hybrid Search Beat 1M-Token Context Stuffing in 2026RAG
RAGVector Search

Beyond Naive RAG: Why Cross-Encoder Reranking & Hybrid Search Beat 1M-Token Context Stuffing in 2026

An architectural guide on production RAG in 2026: why 3-stage hybrid retrieval with cross-encoder rerankers outperforms brute-force 1M-token context stuffing in accuracy, latency, and cost.

Aug 21, 20265 min read
Valkey vs. Redis 8: The 2026 In-Memory Database Benchmark & Architectural GuideValkey
ValkeyRedis

Valkey vs. Redis 8: The 2026 In-Memory Database Benchmark & Architectural Guide

An in-depth engineering benchmark and architectural comparison of Valkey and Redis 8 in 2026: multi-threaded engine performance, RESP3 wire compatibility, and cloud migration strategies.

Aug 20, 20265 min read