Implicit Bias of Gradient Descent: How Optimization Geometry Replaces Explicit Regularization
title: "Implicit Bias of Gradient Descent: How Optimization Geometry Replaces Explicit Regularization" feature_image: "https://cms.llms.blog/content/images/2026/08/implicit-bias-cover.png" status: "published" When engineers train a neural network with plain stochastic gradient descent and no weight decay, the result often generalizes instead of collapsing into an overfit mess. Classical learning theory predicts disaster: with more parameters than data points, unregularized training should find
1 min
