Activation Checkpointing and Selective Rematerialization: Mathematical Foundations, Sub-Linear Memory Scaling, Recomputation Trade-Offs, and Megatron-LM Attention Partitioning
Activation Checkpointing and Selective Rematerialization: Mathematical Foundations, Sub-Linear Memory Scaling, Recomputation Trade-Offs, and Megatron-LM Attention Partitioning Training modern large language models requires orchestrating hundreds of billions of parameters across distributed GPU clusters. While parameter counts and optimizer states are fixed for a given model architecture, the activation memory generated during forward passes scales directly with sequence length, batch size, and



