Knowledge Distillation: Mathematical Foundations, Dark Knowledge, Soft Target Regularization, and Sequence-Level Policy Transfer
Knowledge Distillation: Mathematical Foundations, Dark Knowledge, Soft Target Regularization, and Sequence-Level Policy Transfer Knowledge distillation is a foundational model compression and transfer technique wherein a compact "student" neural network is trained to reproduce the functional behavior, internal representations, or output distributions of a larger, high-capacity "teacher" model or ensemble. First formalized in modern deep learning by Hinton, Vinyals, and Dean (2015), following ea



