Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
Current multimodal agents still fail at precise visual actions tied to what they see.
The paper introduces EASEL, a benchmark where agents paint or mark a canvas from visual evidence, testing closed-loop tool use rather than static answering. Across 25 models, reconstruction scores stayed low, with many runs peaking early and then degrading. A 440k-sample training set and EASEL-9B improved results by 6.3% over its base model, enough to rank third in the evaluation. HF Daily Papers' note
The paper introduces EASEL, a benchmark where agents paint or mark a canvas from visual evidence, testing closed-loop tool use rather than static answering. Across 25 models, reconstruction scores stayed low, with many runs peaking early and then degrading. A 440k-sample training set and EASEL-9B improved results by 6.3% over its base model, enough to rank third in the evaluation. HF Daily Papers' note
score 5