Megadose AI progress, ranked daily.

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

· HF Daily Papers ·
CAFE trains the search agent and its feedback critic in alternation, instead of letting either side stand still.

The paper argues that final-answer rewards miss the point where a search trajectory starts going wrong. CAFE uses a shared-parameter model as both agent and critic, learning when to ask for feedback and how to recover from it. In tests across seven agentic search benchmarks, the authors report better average performance than evaluated RL-based search agents, gains across six out-of-domain benchmarks, and fewer answer-level hallucinations. Ablations found that updating only the agent or only the critic plateaued, while alternating both kept improving. HF Daily Papers' note

score 5

Categories: Research