Disaggregated Prefill and Decode in Production: Architecture, Economics, and KV Transfer Protocols

In production large language model serving, the fundamental architectural tension lies between two computationally distinct phases: prefill (processing the input prompt) and decode (generating output tokens autoregressively). In standard co-located serving systems, both phases share the same GPU instances, memory pools, and execution batches. This co-location creates head-of-line blocking, degrades Time to First Token (TTFT), inflates Time Per Output Token (TPOT), and limits cluster-wide hardwar

5 min
Disaggregated Prefill and Decode in Production: Architecture, Economics, and KV Transfer Protocols

In production large language model serving, the fundamental architectural tension lies between two computationally distinct phases: prefill (processing the input prompt) and decode (generating output tokens autoregressively). In standard co-located serving systems, both phases share the same GPU instances, memory pools, and execution batches. This co-location creates head-of-line blocking, degrades Time to First Token (TTFT), inflates Time Per Output Token (TPOT), and limits cluster-wide hardware utilization.

To resolve this bottleneck, engineering teams and research labs have shifted toward prefill-decode disaggregation. By decoupling prefill and decode workloads across specialized GPU pools connected over high-bandwidth networks, disaggregated serving architectures allow independent scaling, phase-specific parallelism strategies, and tighter latency Service Level Objectives (SLOs).

Disaggregated LLM Serving Architecture

The Compute-Memory Dichotomy

The necessity of disaggregation stems from the contrasting hardware bottlenecks of transformer inference:

  • Prefill is compute-bound: When processing an input prompt of length N, attention and feed-forward operations execute parallel matrix-matrix multiplications (GEMMs). Arithmetic intensity is high, fully saturating GPU Tensor Cores. Prefill operations achieve high compute efficiency even with small batch sizes.
  • Decode is memory-bandwidth bound: Generating tokens one at a time requires loading all model weights and the accumulating Key-Value (KV) cache from high-bandwidth memory (HBM) to on-chip SRAM for each single token. These matrix-vector operations (GEMVs) exhibit low arithmetic intensity. Decode throughput is constrained directly by HBM bandwidth rather than raw FLOPS.

When an inference engine combines both phases into a single batch, such as via standard continuous batching, a long prefill request monopolizes the GPU compute cores. Ongoing decode requests are stalled, causing erratic latency spikes in token delivery.

Limitations of Chunked Prefill

Early mitigations like Sarathi introduced chunked prefill, dividing long prompt computations into smaller token chunks (for example, 512 tokens) and piggybacking them alongside decode tokens in the same forward pass.

While chunked prefill smooths out catastrophic execution stalls, it remains a compromise. Chunking introduces scheduler complexity, increases total kernel launch overhead, and forces prefill and decode to share identical tensor parallel (TP) configurations. A cluster tuned for high-throughput decoding may require high tensor parallelism to aggregate HBM bandwidth across GPUs, whereas prefill workloads often perform more efficiently with pipeline parallelism (PP) or lower tensor parallel degrees to minimize communication overhead.

Architectural Blueprint of Disaggregated Inference

Disaggregated serving separates the physical infrastructure into two distinct tiers governed by a centralized coordinator:

  • Prefill Pool (P-Nodes): Optimized for compute density and large GEMM execution. Requests enter the prefill cluster, generate their initial token, and construct the complete KV cache for the prompt context.
  • Decode Pool (D-Nodes): Optimized for HBM capacity, memory bandwidth, and high-concurrency batching. D-nodes receive the pre-populated KV cache and run continuous decoding loops until completion.
  • KV Cache Transfer Fabric: A low-latency networking layer responsible for transmitting populated KV blocks from P-nodes to D-nodes.

KV Transfer Network Math and Bandwidth Constraints

The primary engineering challenge in disaggregated serving is KV cache migration latency. The size of the KV cache generated during prefill depends on model dimensions, precision, and context length:

KV Cache Size = 2 * L * H_KV * D_head * S_prompt * P_bytes

Where:

  • L is the number of transformer layers.
  • H_KV is the number of Key-Value attention heads.
  • D_head is the head dimension.
  • S_prompt is the input prompt length in tokens.
  • P_bytes is bytes per parameter (2 for FP16/BF16, 1 for FP8).

For a 70-billion parameter model utilizing Grouped-Query Attention (such as Llama 3.1 70B with 80 layers, 8 KV heads, and 128 head dimension):

  • In 16-bit precision, each token requires 327,680 bytes (320 KB) of KV cache across all layers.
  • A 4,096-token prompt generates 1.31 GB of KV cache.
  • An 8,192-token prompt generates 2.62 GB of KV cache.

If transmitted over a standard 100 Gbps network link (theoretical maximum ~12.5 GB/s, practical throughput ~10-11 GB/s), transferring 1.31 GB takes approximately 120 milliseconds. On unoptimized networks, this transfer delay can exceed the prefill compute time itself, negating the latency gains of disaggregation.

Consequently, production disaggregation systems rely on RDMA (Remote Direct Memory Access) over InfiniBand or RoCEv2, alongside layer-wise pipelined streaming. Instead of waiting for the entire forward pass to complete, P-nodes stream early transformer layer KV blocks to D-nodes while deeper layers are still computing.

Production Implementations and System Designs

Several production platforms and research frameworks have established standard patterns for disaggregation:

  • DistServe (OSDI 2024): Developed by researchers at UC San Diego and Peking University, DistServe explicitly decouples TTFT and TPOT objectives. By assigning distinct parallelism plans to each phase and optimizing request placement against cluster bandwidth, DistServe demonstrated up to 7.4 times higher request capacity under strict SLO constraints compared to unified engines.
  • Mooncake (FAST 2025): The serving architecture powering Moonshot AI's Kimi chatbot, Mooncake, implements a KV cache-centric disaggregated design. Mooncake pools underutilized host DRAM, local SSDs, and GPU memory across the cluster into a shared storage layer. Under real-world chatbot traffic with long contexts, Mooncake achieved up to 525% throughput improvements while respecting latency budgets.
  • Together AI CPD: Together AI deployed Cache-Aware Prefill-Decode Disaggregation, segregating traffic into cold and warm tiers based on prefix cache hit rates. By isolating cold prefill requests and reading warm KV blocks directly from distributed caches, the system reported up to 40% higher sustainable throughput on long-context workloads.
  • vLLM and SGLang Connectors: Modern open-source serving engines integrate disaggregated transfer connectors. vLLM implements NIXLConnector and LMCache, enabling out-of-the-box producer-consumer KV transfers over UCX, libfabric, and AWS EFA interfaces.

Engineering Decision Matrix

Disaggregation is not universally optimal for every inference deployment. System architects evaluate the trade-offs based on workload characteristics:

  • Adopt Disaggregated Serving When:
  • Average prompt lengths exceed 2,048 tokens (document analysis, codebases, multi-turn agent histories).
  • Strict, independent SLOs are enforced for TTFT (e.g., <500ms) and TPOT (e.g., <25ms per token).
  • High-speed RDMA networking (200G to 800G InfiniBand/RoCE) or PCIe Gen5 fabrics are available between host nodes.
  • Traffic volume is sufficient to keep dedicated P-node and D-node pools continuously saturated.
  • Retain Co-Located Serving When:
  • Prompts are short (e.g., <512 tokens), where KV transfer overhead outweighs prefill compute savings.
  • Network interconnects are restricted to standard 10 Gbps or 25 Gbps Ethernet without RDMA support.
  • Serving small or low-traffic deployments where provisioning separate pools leads to resource fragmentation and idle GPU hours.

As context windows expand and multi-agent systems generate deeper conversation trees, disaggregated prefill-decode infrastructure is becoming the standard deployment model for hyperscale AI serving platforms.

Sources

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min