Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
AudioRubrics trains audio reasoning models with per-sample reward criteria generated from the waveform and refreshed around the model’s own failures.
The paper argues that final-answer rewards can let models bypass the audio, while fixed process rewards go stale as training improves. Its framework builds audio-grounded rubrics for each sample, then regenerates and reweights them from rollout groups to keep the reward signal targeted. Across three audio reasoning benchmarks, the authors report gains over open-source and training-based baselines. They also say the method settles into stable reasoning lengths instead of collapsing or growing without bound. HF Daily Papers' note
The paper argues that final-answer rewards can let models bypass the audio, while fixed process rewards go stale as training improves. Its framework builds audio-grounded rubrics for each sample, then regenerates and reweights them from rollout groups to keep the reward signal targeted. Across three audio reasoning benchmarks, the authors report gains over open-source and training-based baselines. They also say the method settles into stable reasoning lengths instead of collapsing or growing without bound. HF Daily Papers' note
score 4