Megadose AI progress, ranked and analyzed.

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

· ArXiv · AI/CL/LG ·
Frontier multimodal agents still miss most anomalies in interactive 3D worlds.

The paper introduces WorldAuditBench, with 213 anomaly tasks across 13 Unreal Engine 5 environments and five anomaly families. It tests five models under a fixed exploration budget, comparing a VLA-then-VLM pipeline with an end-to-end VLM agent. Reported success rates run from 6.6% to 42.3%, far below human performance at 83.4%. The benchmark is meant to probe whether agents can use visual reasoning to guide action and evidence-gathering in 3D spaces. ArXiv · AI/CL/LG's note

score 4

Categories: Research