State Space Models and Mamba (Mamba-1 and Mamba-2): Mathematical Foundations, Selective State Spaces, Structured State Space Duality (SSD), and Linear-Time Sequence Modeling
State Space Models (SSMs) and their modern selective formulations, most notably Mamba-1 and Mamba-2, represent a foundational alternative to the standard Transformer architecture for sequence modeling. While multi-head self-attention scales quadratically with sequence length ($O(T^2)$) and requires an ever-expanding Key-Value (KV) cache during autoregressive generation ($O(T)$), State Space Models achieve linear time complexity ($O(T)$) during training and constant memory footprint ($O(1)$) per

