CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
CAFE trains a search agent and its feedback critic in alternating roles so the corrections change as the agent improves.
The paper argues that final-answer rewards are too blunt for search agents because they miss errors that happen mid-trajectory. CAFE uses a shared-parameter model as both agent and critic, learning when to request feedback and how to recover from it. In tests across seven agentic search benchmarks, it beat the evaluated RL-based search agents on average and held gains across six out-of-domain benchmarks. The authors report fewer answer-level hallucinations, and ablations showed that updating only the agent or only the critic eventually stalled. ArXiv · AI/CL/LG's note
The paper argues that final-answer rewards are too blunt for search agents because they miss errors that happen mid-trajectory. CAFE uses a shared-parameter model as both agent and critic, learning when to request feedback and how to recover from it. In tests across seven agentic search benchmarks, it beat the evaluated RL-based search agents on average and held gains across six out-of-domain benchmarks. The authors report fewer answer-level hallucinations, and ablations showed that updating only the agent or only the critic eventually stalled. ArXiv · AI/CL/LG's note
score 5