KV Cache Compression and Dynamic Pruning in Production Serving: Comparing StreamingLLM, H2O, SnapKV, PyramidKV, and DuoAttention Architecture, Memory Bandwidth Reduction, and Needle Retrieval Retention
KV Cache Compression and Dynamic Pruning in Production Serving: Comparing StreamingLLM, H2O, SnapKV, PyramidKV, and DuoAttention Architecture, Memory Bandwidth Reduction, and Needle Retrieval Retention As context windows in production large language models expand from 8K tokens to 128K tokens and beyond, the Key-Value (KV) cache replaces model weights as the primary consumer of GPU High Bandwidth Memory (HBM). In high-throughput serving systems, the KV cache grows linearly with sequence length


