Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
Early context pruning delivered the biggest token savings, with heuristics cutting usage by up to 73% with little quality loss.
The paper studies when deep research agents should stop carrying extra evidence through retrieval and synthesis. Its comparison finds the pruning stage matters more than the exact scoring rule. Early pruning produces the strongest end-to-end efficiency gains, while later pruning mostly cleans up the final synthesis context. Learned pruning is competitive in some trade-offs, but no method wins across quality, efficiency, and faithfulness. HF Daily Papers' note
The paper studies when deep research agents should stop carrying extra evidence through retrieval and synthesis. Its comparison finds the pruning stage matters more than the exact scoring rule. Early pruning produces the strongest end-to-end efficiency gains, while later pruning mostly cleans up the final synthesis context. Learned pruning is competitive in some trade-offs, but no method wins across quality, efficiency, and faithfulness. HF Daily Papers' note
score 5