Megadose AI progress, ranked and analyzed.

ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

· ArXiv · AI/CL/LG ·
The paper says exponential-noise BoN gives finer reward-control at inference time and converges fast under exact finite-budget analysis.

ExpBoN replaces hard best-of-$n$ selection with an exponential-noise report-noisy-max mechanism, giving a soft alternative with smoother control over reward versus distribution shift. The authors prove exponentially fast convergence in total variation, expected reward, and both KL directions. They also plug it into guided speculative inference as ExpGSI, cutting estimated compute while keeping comparable accuracy. Reported experiments show compute reductions of 14%–39% for Qwen2.5-Math and up to 45% for Qwen3 at $n=16$ on the tested math and STEM benchmarks. ArXiv · AI/CL/LG's note

score 6

Categories: Research