026

THE METER

SECTION 03

ISSUE 001

OBSERVEDinfrastructurenowconfidence / high

Prefill and Decode Split Apart

Observed now: NVIDIA separately scales prefill and decode pools and transfers KV state between them; Google’s TPU 8i is purpose-built for inference and reinforcement learning with larger on-chip SRAM, HBM, and lower-latency collectives. Serving is splitting along workload physics. Superior production economics still require measurement on real traffic.

Why this idea is here

What the evidence establishes.

NVIDIA documents separately scalable prefill and decode pools, KV-state transfer, cache-aware routing, and distributed inference; Google Cloud documents TPU 8i as its eighth-generation serving specialist and TPU 8t as its training system. Together they support specialization across serving stages and silicon roles.

Source ledger

Read the sources.

  1. S01
    NVIDIA Dynamo: Disaggregated Serving

    official technical documentation / published 2026-07-09 / retrieved 2026-07-09

  2. S02
    TPU 8t, TPU 8i, and Virgo for the agentic era

    official infrastructure announcement / published 2026-04-22 / retrieved 2026-07-10

Back to all 500 ideas