Llm-Inference12/21 - Cache the Right Thing: Prompt, Semantic, and Cost-Aware ReuseJuly 12, 20262/21 - Smaller Numbers, Faster Models: Quantization and BatchingJuly 2, 202620/20 - Expert Parallelism: Routing Tokens Through a City of SpecialistsJune 30, 202616/20 - Streaming Generation: The First Token Is a Product DecisionJune 25, 202615/20 - Memory Offloading: Trading Bandwidth for CapacityJune 24, 202610/20 - Tensor Parallelism: Splitting One Layer Across Many GPUsJune 19, 20269/20 - Quantized Kernels: Why a 4-Bit Model Is Not Automatically FastJune 18, 20267/20 - Parallel Decoding: Predicting More Than One Future at a TimeJune 16, 20266/20 - Early Exit Decoding: Stop Computing Once the Answer Is ClearJune 15, 20263/20 - FlashAttention: Why Moving Fewer Bytes Beats Doing Fewer FLOPsJune 12, 20262/20 - Speculative Decoding: Let a Small Model Guess, Let a Large Model JudgeJune 11, 20261/20 - KV Caching: The Memory That Makes Token Generation PossibleJune 10, 2026