Temperature Scaling and Model Calibration in Deep Neural Networks: Mathematical Foundations, Expected Calibration Error, and Post-Hoc Logit Optimization
Modern deep neural networks achieve high classification accuracy and generative benchmark performance across vision, language, and decision-making tasks. However, optimization for raw accuracy does not ensure that predicted softmax probabilities correspond to true posterior probabilities. A model that assigns a 0.90 probability to an output should be correct exactly 90% of the time. When empirical accuracy systematically diverges from predicted confidence, the model is miscalibrated. Research b
1 min
