New Technical Chapter · CXMT · Memory Wall · July 27, 2026
A dedicated 14-step chapter on CXMT's shipping DDR5 and LPDDR5X products, DRAM cell physics, process and yield, the reported HBM3 path, TSVs, thinning, stacking, bonding, the logic base die, packaging, accelerator qualification, useful bandwidth, and accepted-task economics. Every field is labeled official, reported, teardown, patent, derived, or unknown.
Open the CXMT chapter →
Systems Guide · HBM · Published July 25 · Updated July 27, 2026
A step-by-step guide to the memory behind AI. Start with a coding-agent request or a generated video, then follow the data through software, the GPU, HBM, the rack, power, cooling, manufacturing, and cost. Includes a guided visual map, a 184-scene evidence archive, a deterministic GLM-5.2 matrix micro-lab, and the new complete CXMT memory-wall chapter.
Read the complete guide →
Computex Recap · AI Factories · KV Cache · June 9, 2026
A short Computex recap on KV cache offload as real hardware: Penguin Solutions' CXL KV cache server, Astera Labs' PCIe 6 fabric, how CXL rides PCIe, and one coding-agent example with compaction, step by step.
Read the recap →
Research Note · Automated CUDA · May 28, 2026
A practical founder-engineer map of full-stack inference optimization: workload replay, CUDA kernels, KV cache, routing, hardware, energy, and cost per successful task.
Read the post →
How-to · Local Inference · May 13, 2026
A step-by-step Mac MLX runbook for one specific problem: showing why model choice, quantization, context shape, tool use, and hardware memory layout all change the right inference optimization. Start on a 16GB M5, then map the evidence to LMCache MP, vLLM/Mooncake, SGLang, P/D disaggregation, and larger Qwen3.6 or Kimi K2.6 stacks.
Read the guide →
Briefing · KV Cache · May 13, 2026
A 50-turn OpenClaw or Hermes agent should not keep paying for turn 1. A briefing on prompt layout, host-side shared KV, distributed lookup, RDMA transfer, encoder reuse, and what InferGuard is trying to prove without overclaiming.
Read the briefing →
Runtime Boundary · TokenSpeed · May 7, 2026
A plain-English guide to TokenSpeed, Blackwell, Dynamo, Vera Rubin, KV cache, prefill/decode, and why agentic inference is changing the CPU/GPU runtime boundary.
Read the briefing →