Megadose AI progress, ranked and analyzed.

SceneActBench: Can Agents Act on the 3D Scenes They See?

· HF Daily Papers ·
SceneActBench tests whether VLM agents can take geometrically correct actions in full 3D scenes, and the evaluated models still fall short.

The benchmark covers five visually conditioned 3D tasks using PNG images or sampled video frames, with 3D assets supplied where needed. Agents act through a fixed environment loop, and final outputs are scored against hidden ground truth with task-specific geometric metrics. It contains 520 task cases from 210 source instances. Across eleven proprietary VLM configurations, overall scores ranged from 38.6 to 50.2, with no model performing consistently well across tasks. Source: HF Daily Papers' note.

score 5

Categories: Research