The Edge of Stability: How Progressive Sharpening and Hessian Dynamics Govern Deep Learning Optimization
In classical convex optimization, the behavior of gradient descent is dictated by the Lipschitz smoothness constant of the objective function. If a function $f(\theta)$ has an $L$-smooth gradient—meaning the largest eigenvalue of its Hessian matrix is bounded by $\lambda_{\max}(\nabla^2 f(\theta)) \le L$—gradient descent with learning rate $\eta$ is guaranteed to monotonically reduce the loss if and only if $\eta < 2/L$. When the step size exceeds this threshold ($\eta > 2/\lambda_{\max}$), stan


