Double descent is the principle of least action
The paper argues double descent follows from a statistical-mechanics view of stochastic-gradient training.
It models training as a particle diffusing across the loss landscape at an induced temperature. Finite training time acts like effective weight decay, turning each parameter into a quadratic degree of freedom. As parameter count rises at fixed training loss, equipartition lowers the temperature and pushes samples toward lower-norm stationary paths. Source: ArXiv · AI/CL/LG's note.
It models training as a particle diffusing across the loss landscape at an induced temperature. Finite training time acts like effective weight decay, turning each parameter into a quadratic degree of freedom. As parameter count rises at fixed training loss, equipartition lowers the temperature and pushes samples toward lower-norm stationary paths. Source: ArXiv · AI/CL/LG's note.
score 4