SoftServe: A Scalable Quasi-Newton Method for Deep Learning
SoftServe makes quasi-Newton training practical for very large, non-convex neural networks.
The paper introduces a family of methods that keeps curvature estimates positive-definite without line searches or ad hoc fixes. It includes diagonal and Kronecker-factored variants meant to scale to massive models. The authors use coupled Newton-Schulz iterations so the needed matrix work can run as GPU-friendly multiplications instead of decompositions. They report lower losses than Adam, Muon, and SOAP on ill-conditioned tasks, including recurrent nets, autoencoders, physics-informed networks, and a 136M-parameter physics-informed diffusion model. ArXiv · AI/CL/LG's note
The paper introduces a family of methods that keeps curvature estimates positive-definite without line searches or ad hoc fixes. It includes diagonal and Kronecker-factored variants meant to scale to massive models. The authors use coupled Newton-Schulz iterations so the needed matrix work can run as GPU-friendly multiplications instead of decompositions. They report lower losses than Adam, Muon, and SOAP on ill-conditioned tasks, including recurrent nets, autoencoders, physics-informed networks, and a 136M-parameter physics-informed diffusion model. ArXiv · AI/CL/LG's note
score 4