Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
Small proxy runs may be enough to set the SFT-versus-RL data split for larger models.
The paper frames annotation budgeting as a near-optimal region rather than a single best ratio. It reports that those regions stay broad even within 2-10% of peak performance, widen with model scale, and transfer from smaller proxy models to larger targets. The result is meant to reduce exhaustive large-model searches across SFT and RL allocations. The authors say the pattern holds across tasks, model families, and both off-policy preference RL and on-policy reward supervision. ArXiv · AI/CL/LG's note
The paper frames annotation budgeting as a near-optimal region rather than a single best ratio. It reports that those regions stay broad even within 2-10% of peak performance, widen with model scale, and transfer from smaller proxy models to larger targets. The result is meant to reduce exhaustive large-model searches across SFT and RL allocations. The authors say the pattern holds across tasks, model families, and both off-policy preference RL and on-policy reward supervision. ArXiv · AI/CL/LG's note
score 4