Touchdown Labs Blog

Articles on AI systems, inference, kernels, GPUs, hardware, materials, and manufacturing.

New Technical Chapter · CXMT · Memory Wall · July 27, 2026

CXMT and the Memory Wall: From One DRAM Cell to a Qualified HBM System

A dedicated 14-step chapter on CXMT's shipping DDR5 and LPDDR5X products, DRAM cell physics, process and yield, the reported HBM3 path, TSVs, thinning, stacking, bonding, the logic base die, packaging, accelerator qualification, useful bandwidth, and accepted-task economics. Every field is labeled official, reported, teardown, patent, derived, or unknown.

Open the CXMT chapter →
Systems Guide · HBM · Published July 25 · Updated July 27, 2026

How HBM Works for AI: From One Memory Cell to a Real AI Task

A step-by-step guide to the memory behind AI. Start with a coding-agent request or a generated video, then follow the data through software, the GPU, HBM, the rack, power, cooling, manufacturing, and cost. Includes a guided visual map, a 184-scene evidence archive, a deterministic GLM-5.2 matrix micro-lab, and the new complete CXMT memory-wall chapter.

Read the complete guide →
Computex Recap · AI Factories · KV Cache · June 9, 2026

Computex 2026: Seeing the Physical Layer of the AI Factory

A short Computex recap on KV cache offload as real hardware: Penguin Solutions' CXL KV cache server, Astera Labs' PCIe 6 fabric, how CXL rides PCIe, and one coding-agent example with compaction, step by step.

Read the recap →
Research Note · Automated CUDA · May 28, 2026

Automated CUDA Is Almost Here. Revenue and Capability per GPU Are Doubling.

A practical founder-engineer map of full-stack inference optimization: workload replay, CUDA kernels, KV cache, routing, hardware, energy, and cost per successful task.

Read the post →
How-to · Local Inference · May 13, 2026

Set Up a Local OpenClaw or Hermes KV-Cache Learning Loop with Mac MLX

A step-by-step Mac MLX runbook for one specific problem: showing why model choice, quantization, context shape, tool use, and hardware memory layout all change the right inference optimization. Start on a 16GB M5, then map the evidence to LMCache MP, vLLM/Mooncake, SGLang, P/D disaggregation, and larger Qwen3.6 or Kimi K2.6 stacks.

Read the guide →
Briefing · KV Cache · May 13, 2026

KV Cache Is Becoming the Memory Hierarchy of Inference

A 50-turn OpenClaw or Hermes agent should not keep paying for turn 1. A briefing on prompt layout, host-side shared KV, distributed lookup, RDMA transfer, encoder reuse, and what InferGuard is trying to prove without overclaiming.

Read the briefing →
Runtime Boundary · TokenSpeed · May 7, 2026

TokenSpeed, Blackwell, and Vera Rubin: The Runtime Boundary Is Moving

A plain-English guide to TokenSpeed, Blackwell, Dynamo, Vera Rubin, KV cache, prefill/decode, and why agentic inference is changing the CPU/GPU runtime boundary.

Read the briefing →