Megadose AI progress, ranked and analyzed.

BiasReducer: Adaptive Bias Mitigation for Reward Models

· HF Daily Papers ·
BiasReducer edits only the reward head, then chooses which bias corrections matter for a new dataset.

The paper says reward models can overvalue surface traits like length or confidence, pushing LLMs toward answers that score better without being more correct. BiasReducer learns which attributes a reward model is sensitive to, estimates how to reduce each dependence, and applies the relevant edits at evaluation time. Across five reward models, its BiasReducer-M variant improved three bias-robustness benchmarks by 8.3, 18.0, and 6.9 percentage points on average. The authors also report downstream reductions in unnecessary verbosity and sycophancy while keeping judged quality comparable. HF Daily Papers' note

score 4

Categories: Research