THE METER
SECTION 03
ISSUE 001
CXL Pools Become Model Memory
Projection: Qualcomm’s roadmap emphasizes near-memory, disaggregated inference; NVIDIA already transfers KV state between independently scaled pools. Composable memory could hold caches, retrieval state, or model components beyond accelerator HBM. The forecast is supported only if pooled capacity outweighs added latency, coherence, privacy, and orchestration costs.
Why this idea is here
What the evidence establishes.
Qualcomm announces agentic CPUs, near-memory high-bandwidth compute, disaggregated inference, liquid-cooled racks, and claimed improvements in bandwidth per watt; NVIDIA documents separately scalable prefill and decode pools, KV-state transfer, cache-aware routing, and distributed inference. These are source-backed premises for this projection; they do not by themselves prove broad adoption or the eventual outcome.
Source ledger
Read the sources.
- S01Qualcomm Dragonfly Data Center Roadmap
official hardware announcement / published 2026-06-24 / retrieved 2026-07-09
- S02NVIDIA Dynamo: Disaggregated Serving
official technical documentation / published 2026-07-09 / retrieved 2026-07-09