KV Cache Eviction and Streaming Attention: Mathematical Foundations, Attention Sink Mechanics, Heavy Hitter Oracles (H2O), and Bounded-Memory Generation
KV Cache Eviction and Streaming Attention: Mathematical Foundations, Attention Sink Mechanics, Heavy Hitter Oracles (H2O), and Bounded-Memory Generation In autoregressive transformer generation, memory consumption and serving throughput are dominated by the Key-Value (KV) cache. For long-context generation and continuous multi-turn dialogue, the linear growth of the KV cache with sequence length imposes an unsustainable memory footprint and saturates high-bandwidth GPU memory channels. Standar



















