Megadose AI progress, ranked and analyzed.

The Loss Does Not See the Basis, but Adam Does

· ArXiv · AI/CL/LG ·
Adam can choose a different interpolating solution from gradient descent even when the loss is unchanged by basis rotations.

The paper traces that split to gauge symmetry in factored models, where the loss is invariant under shared rotations of the factors. It says gradient descent and several equivariant optimizers preserve the low-rank bias of gradient flow, while Adam, RMSProp, and other coordinate-wise methods do not. Experiments sort nine update rules on matrix sensing and report that moving from coordinate-wise to shared-scalar preconditioning restores the bias monotonically. The author also reports transformer and hyperspectral examples where the optimizer’s basis sensitivity shows up in invariants and held-out error. ArXiv · AI/CL/LG's note

score 5

Categories: Research