Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Beacon is built to make visual-reasoning tools conditional, not automatic.
The paper argues that current agentic visual reasoning models often fail at deciding when tools are actually needed. Its analysis says tool gains on hard examples are largely canceled by mistakes introduced on easier ones. Beacon adds reinforcement-learning mechanisms meant to reward necessary tool use and improve tool handling on the hardest cases. The authors report stronger overall results across multiple benchmarks, with better Mode Adaptiveness and Tool Effect. HF Daily Papers' note
The paper argues that current agentic visual reasoning models often fail at deciding when tools are actually needed. Its analysis says tool gains on hard examples are largely canceled by mistakes introduced on easier ones. Beacon adds reinforcement-learning mechanisms meant to reward necessary tool use and improve tool handling on the hardest cases. The authors report stronger overall results across multiple benchmarks, with better Mode Adaptiveness and Tool Effect. HF Daily Papers' note
score 5