What Is a KV Cache? The Memory Bottleneck Behind LLM Serving

What Is a KV Cache? The Memory Bottleneck Behind LLM Serving Every token a large language model generates costs more than compute. It costs memory. The mechanism that stores the model's working state across a generation is called the KV cache, and managing it well is what separates a serving system that feels fast from one that stalls. This explainer covers what the KV cache is, why it wastes memory, and how the PagedAttention algorithm from Berkeley's vLLM project reshaped how it is handled.

4 min
What Is a KV Cache? The Memory Bottleneck Behind LLM Serving

What Is a KV Cache? The Memory Bottleneck Behind LLM Serving

Every token a large language model generates costs more than compute. It costs memory. The mechanism that stores the model's working state across a generation is called the KV cache, and managing it well is what separates a serving system that feels fast from one that stalls. This explainer covers what the KV cache is, why it wastes memory, and how the PagedAttention algorithm from Berkeley's vLLM project reshaped how it is handled.

What the KV cache stores

When a transformer generates text, it produces tokens one at a time. For each token in the input, the model computes attention keys and values, the tensors it uses to decide how much each earlier token should influence the next one. Instead of recomputing those keys and values for every token on every step, serving systems cache them in GPU memory and reuse them as each new token arrives.

These cached tensors are called the key-value cache, or KV cache. The vLLM team, which built the PagedAttention algorithm, describes the KV cache as both large and dynamic: a single sequence in a LLaMA-13B model can use up to 1.7 GB of cache, and the size grows with sequence length, which is highly variable across requests.

Why it is the serving bottleneck

In the autoregressive decoding loop, the model generates output one token at a time, and each step depends on all the tokens before it. That sequential dependency means a serving system cannot hide latency by computing tokens in parallel. The practical lever for throughput is batching: running many requests through the GPU at once so the expensive matrix multiplications are amortized. But every request consumes KV cache memory for its whole lifetime, so the amount of memory available effectively caps how many requests can be batched.

The vLLM team found that existing serving systems handled this poorly. As they wrote, "existing systems waste 60-80% of memory due to fragmentation and over-reservation." When a request finishes, its cache allocation leaves a gap the system cannot always reuse for a different request of a different length. The system must also reserve enough space for the longest possible sequence, even when most requests are short.

PagedAttention: memory as a page table

PagedAttention is the attention algorithm that vLLM introduced to solve this. It borrows the concept of virtual memory and paging from operating systems. Instead of allocating one contiguous block of memory per sequence, PagedAttention splits each sequence's KV cache into fixed-size blocks, and it is fine for those blocks to live in non-contiguous physical memory.

PagedAttention partitions a sequence's KV cache into fixed-size, non-contiguous blocks, reducing fragmentation

A block table maps a sequence's logical blocks to their physical locations, in the same way an operating system maps a process's virtual pages to physical frames. Physical blocks are allocated on demand, only as new tokens are generated. Because the blocks need not be contiguous, the gaps left by finished requests can be reused by other sequences, which nearly eliminates fragmentation. The vLLM team reports that memory waste falls to under 4%, limited mainly to the last, partially filled block of a sequence.

This memory efficiency has a direct consequence for throughput: with less memory wasted, more sequences can be batched onto the GPU at once. In its evaluations, vLLM reported 14-24x higher throughput than the HuggingFace Transformers library and 2.2-3.5x higher than HuggingFace's Text Generation Inference, with no change to model architecture. The accompanying paper reports 2-4x throughput gains versus earlier systems such as FasterTransformer and Orca at the same latency, with larger gains on longer sequences and bigger models.

Sharing the cache across outputs

PagedAttention also enables a second optimization: memory sharing. Many sampling strategies, such as parallel sampling and beam search, generate several output candidates from the same prompt. Those candidates all reuse the prompt's KV cache, because the prompt tokens are identical.

PagedAttention lets different sequences map their logical blocks to the same physical blocks, with reference counting and a copy-on-write mechanism to keep the data safe when a candidate diverges. The vLLM team reports this sharing cuts memory use for these sampling algorithms by up to 55%, which can translate to up to 2.2x higher throughput.

Why this matters

KV cache management sits underneath most of the cost and latency that users feel in LLM APIs and local deployments. It is why the same model can feel dramatically faster or slower depending on the serving software, and it is a large part of why inference costs fell so fast in the years after the original transformer architectures became standard.

PagedAttention is now the foundation of vLLM, one of the most widely deployed open-source inference engines, and versions of block-based KV cache management appear throughout the serving ecosystem. For anyone working with LLMs operationally, the KV cache is the concept that explains why serving performance is about far more than raw model size or GPU count.

Sources

Written by

More to read

  • Modular Open-Sources Mojo Language Compiler and Toolchain Under Apache 2.0

    Modular Open-Sources Mojo Language Compiler and Toolchain Under Apache 2.0 Modular has released the complete source code for the Mojo programming language compiler, standard tooling, and runtime infrastructure under the Apache 2.0 license with LLVM exceptions. The announcement, delivered on August 18, 2026 during the company's ModCon developer conference, fulfills a multi-year roadmap commitment to transition the systems programming language to a fully open development model. The compiler sour

    1 min
  • AI Agent Evaluation in Production: Trajectory Benchmarks, Sandbox Harnesses, and Flakiness Mitigation

    Evaluating standard large language models relies on static input-output pairs: a fixed prompt produces a completion that an automated script compares against reference strings or grades with a calibrated judge. Autonomous AI agents break this paradigm completely. An agent executes a multi-step trajectory consisting of planning, tool invocation, environment state observation, error recovery, and variable-length decision loops. Evaluating an agent requires testing not just the final string output,

    1 min
  • Fully Sharded Data Parallel (FSDP) and ZeRO: How Memory Sharding Eliminates Redundant Model States in Distributed Training

    Fully Sharded Data Parallel (FSDP) and ZeRO: How Memory Sharding Eliminates Redundant Model States in Distributed Training Training large language models across distributed GPU clusters introduces a fundamental memory bottleneck. In traditional Distributed Data Parallel (DDP) setups, every GPU maintains an identical copy of model weights, optimizer states, and gradients while processing independent data batches. As models scale from billions to hundreds of billions of parameters, static model s

    1 min