The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
The paper argues that crop-and-zoom often changes the process more than the answer.
The authors test whether visual evidence returned by active image tools actually causes multimodal models to answer better. Across six models and five fine-grained perception benchmarks, they find many runs where tool calls add token cost but little or no causal benefit. They name two failure modes: models call tools without using the returned evidence, or get useful observations but schedule calls incoherently. HF Daily Papers' note
The authors test whether visual evidence returned by active image tools actually causes multimodal models to answer better. Across six models and five fine-grained perception benchmarks, they find many runs where tool calls add token cost but little or no causal benefit. They name two failure modes: models call tools without using the returned evidence, or get useful observations but schedule calls incoherently. HF Daily Papers' note
score 5