Probe-Space Preconditioning for Fast and Stable Zero-Order Training
A zero-order training variant claims backprop-level gains while keeping memory near inference mode.
The paper says 1.5-SPSA adds one clean forward pass to 1SPSA to build a cheap probe-space diagonal preconditioner. That lets it down-weight high-curvature directions and converge faster than prior zero-order solvers in the authors’ tests. On six post-training datasets across Qwen3 and OPT models, they report state-of-the-art zero-order results with far fewer optimization steps. One highlighted run trains OPT-13B to beat MeZO and BP on SST-2 by 3.1 percentage points in 70 steps, versus MeZO’s 100,000. ArXiv · AI/CL/LG's note
The paper says 1.5-SPSA adds one clean forward pass to 1SPSA to build a cheap probe-space diagonal preconditioner. That lets it down-weight high-curvature directions and converge faster than prior zero-order solvers in the authors’ tests. On six post-training datasets across Qwen3 and OPT models, they report state-of-the-art zero-order results with far fewer optimization steps. One highlighted run trains OPT-13B to beat MeZO and BP on SST-2 by 3.1 percentage points in 70 steps, versus MeZO’s 100,000. ArXiv · AI/CL/LG's note
score 5