THE METER
SECTION 03
ISSUE 001
Prefill and Decode Split Apart
Observed now: NVIDIA separately scales prefill and decode pools and transfers KV state between them; Google’s TPU 8i is purpose-built for inference and reinforcement learning with larger on-chip SRAM, HBM, and lower-latency collectives. Serving is splitting along workload physics. Superior production economics still require measurement on real traffic.
Why this idea is here
What the evidence establishes.
NVIDIA documents separately scalable prefill and decode pools, KV-state transfer, cache-aware routing, and distributed inference; Google Cloud documents TPU 8i as its eighth-generation serving specialist and TPU 8t as its training system. Together they support specialization across serving stages and silicon roles.
Source ledger
Read the sources.
- S01NVIDIA Dynamo: Disaggregated Serving
official technical documentation / published 2026-07-09 / retrieved 2026-07-09
- S02TPU 8t, TPU 8i, and Virgo for the agentic era
official infrastructure announcement / published 2026-04-22 / retrieved 2026-07-10