Chinchilla Scaling Laws: Mathematical Foundations of Compute-Optimal Pre-Training, IsoFLOP Loss Profiles, Parametric Power Laws, and Data-Compute Allocation
Chinchilla Scaling Laws: Mathematical Foundations of Compute-Optimal Pre-Training, IsoFLOP Loss Profiles, Parametric Power Laws, and Data-Compute Allocation When allocating a fixed computational budget to train an autoregressive Transformer, engineers face a fundamental trade-off: should FLOPs be spent increasing the model parameter count ($N$), or should they be spent streaming a larger volume of training tokens ($D$)? For several years, frontier AI development followed the empirical scaling


