ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
The paper’s claim is that step-level rewards can make a 4B search agent competitive with much larger systems.
ABSeeker is trained with Answer-Backtracked Credit Assignment, which traces from the known answer to intermediate clues and scores each search step against them. The method rewards useful steps even in failed trajectories and downweights erroneous or redundant ones during SFT and GRPO. Built on Qwen3.5-4B with 8.5k examples, it reports 37.3% on BrowseComp and 39.1% on BrowseComp-ZH, rising to 55.3% and 52.9% with context management. ArXiv · AI/CL/LG's note
ABSeeker is trained with Answer-Backtracked Credit Assignment, which traces from the known answer to intermediate clues and scores each search step against them. The method rewards useful steps even in failed trajectories and downweights erroneous or redundant ones during SFT and GRPO. Built on Qwen3.5-4B with 8.5k examples, it reports 37.3% on BrowseComp and 39.1% on BrowseComp-ZH, rising to 55.3% and 52.9% with context management. ArXiv · AI/CL/LG's note
score 5