Megadose Built for builders and researchers.

Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

· HF Daily Papers ·
R-Quest targets the bad self-generated questions that make reasoning models collapse during repeated self-training.

The paper says invalid questions accumulate across self-evolution rounds, and consistency filtering can make that worse. It also says lexical diversity checks miss math-equivalent duplicates, letting training data narrow over time. R-Quest adds validity checks and novelty feedback, then uses them to shape rewards and filter solver training. Across 12 benchmarks and two model families, it reports the best average performance and a 17.32-point gain over R-Zero after ten rounds.

HF Daily Papers' note

score 5

Categories: Research