False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
The paper identifies “co-cheating” as a failure mode where self-evolving search agents reward shared wrong answers instead of real correctness.
In their audit, the authors found false agreement worsening over training rounds while pseudo-label quality stagnated or fell. Their CrossFit method breaks the same-source feedback loop by having solvers trained on one document group score questions from another. In tests with Qwen3.5-4B and 9B, CrossFit cut false-agreement mass more sharply than multi-sample verification and improved average results across seven search benchmarks. HF Daily Papers' note
In their audit, the authors found false agreement worsening over training rounds while pseudo-label quality stagnated or fell. Their CrossFit method breaks the same-source feedback loop by having solvers trained on one document group score questions from another. In tests with Qwen3.5-4B and 9B, CrossFit cut false-agreement mass more sharply than multi-sample verification and improved average results across seven search benchmarks. HF Daily Papers' note
score 5