VISTA: A Visual Harness for Reasoning in an Interactive World
VISTA gives Claude Opus 5.0 a visual memory and retrieval loop that takes it to a perfect ARC-AGI-3 public-games score.
The harness lets a multimodal model keep past visual observations losslessly and pull them back into view while reasoning. In the paper’s ARC-AGI-3 run, Claude Opus 5.0 rises from 40.68 to 100.00 Relative Human Action Efficiency and completes all 25 public games. The authors say it used 57.4% fewer actions than first-time human participants. They also report stronger results than minimal-harness baselines across three other visual games and puzzle benchmarks. ArXiv · AI/CL/LG's note
The harness lets a multimodal model keep past visual observations losslessly and pull them back into view while reasoning. In the paper’s ARC-AGI-3 run, Claude Opus 5.0 rises from 40.68 to 100.00 Relative Human Action Efficiency and completes all 25 public games. The authors say it used 57.4% fewer actions than first-time human participants. They also report stronger results than minimal-harness baselines across three other visual games and puzzle benchmarks. ArXiv · AI/CL/LG's note
score 6