VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning
The paper’s core claim is that VLM agents need a managed visual evidence state, not an ever-growing image context.
VLM-in-Sandbox keeps generated crops, masks, overlays, zooms, and renderings in an image ledger while limiting what stays active for the model. The model can promote selected visual evidence for later inspection, separating tool-generated evidence from context management. In tests across seven benchmarks and four base VLMs, the framework reported the highest sample-weighted average accuracy against the listed baselines. A matched study found 66.27% accuracy with 18.6% fewer total tokens than a retain-all control. ArXiv · AI/CL/LG's note
VLM-in-Sandbox keeps generated crops, masks, overlays, zooms, and renderings in an image ledger while limiting what stays active for the model. The model can promote selected visual evidence for later inspection, separating tool-generated evidence from context management. In tests across seven benchmarks and four base VLMs, the framework reported the highest sample-weighted average accuracy against the listed baselines. A matched study found 66.27% accuracy with 18.6% fewer total tokens than a retain-all control. ArXiv · AI/CL/LG's note
score 5