Skip to content

Recent Articles

85 posts · sorted by date
July 16, 2026 10 min · read

16/21 - Agents Need Infrastructure Too: MCP and Workflow Orchestration

A production-focused guide to agents need infrastructure too: mcp and workflow orchestration, with architecture, capacity math, failure analysis, and operational controls.

July 15, 2026 10 min · read

15/21 - The Router Is Part of the Model: Routing, Hedging, and Fallback

A production-focused guide to the router is part of the model: routing, hedging, and fallback, with architecture, capacity math, failure analysis, and operational controls.

July 14, 2026 10 min · read

14/21 - Benchmarking Without Lying: Evals, Load Tests, and A/B Experiments

A production-focused guide to benchmarking without lying: evals, load tests, and a/b experiments, with architecture, capacity math, failure analysis, and operational controls.

July 13, 2026 10 min · read

13/21 - Can You Debug a Token? Observability for AI Systems

A production-focused guide to can you debug a token? observability for ai systems, with architecture, capacity math, failure analysis, and operational controls.

July 12, 2026 10 min · read

12/21 - Cache the Right Thing: Prompt, Semantic, and Cost-Aware Reuse

A production-focused guide to cache the right thing: prompt, semantic, and cost-aware reuse, with architecture, capacity math, failure analysis, and operational controls.

July 11, 2026 10 min · read

11/21 - RAG That Survives Production: Embeddings, Retrieval, and Evidence

A production-focused guide to rag that survives production: embeddings, retrieval, and evidence, with architecture, capacity math, failure analysis, and operational controls.

July 10, 2026 10 min · read

10/21 - Fine-Tuning Without the Full Bill: LoRA, QLoRA, and PEFT

A production-focused guide to fine-tuning without the full bill: lora, qlora, and peft, with architecture, capacity math, failure analysis, and operational controls.

July 9, 2026 10 min · read

9/21 - One Model, Many Accelerators: Multi-GPU and Multi-Node Inference

A production-focused guide to one model, many accelerators: multi-gpu and multi-node inference, with architecture, capacity math, failure analysis, and operational controls.

July 8, 2026 10 min · read

8/21 - The Fabric Between GPUs: NCCL, InfiniBand, RoCE, and GPUDirect

A production-focused guide to the fabric between gpus: nccl, infiniband, roce, and gpudirect, with architecture, capacity math, failure analysis, and operational controls.

July 7, 2026 10 min · read

7/21 - Kubernetes Meets GPUs: Containers, Scheduling, and Isolation

A production-focused guide to kubernetes meets gpus: containers, scheduling, and isolation, with architecture, capacity math, failure analysis, and operational controls.