Megadose AI progress, ranked and analyzed.

SceneActBench: Can Agents Act on the 3D Scenes They See?

· ArXiv · AI/CL/LG ·
SceneActBench finds current VLM agents still weak at acting reliably inside full 3D scenes.

The benchmark tests visually conditioned action across five 3D tasks, using images or sampled video frames and sometimes supplied 3D assets. Each agent runs through the same agent-environment loop, with final outputs checked against hidden ground truth using geometric metrics. The paper reports 520 task cases from 210 source instances. Across eleven proprietary VLM configurations, overall scores range from 38.6 to 50.2, with no model consistently strong across tasks. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research