STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
A new benchmark tests whether VLMs actually understand social navigation scenes well enough for robots to use them.
The paper introduces SocialNav-SUB, a VQA benchmark for spatial, spatiotemporal, and social reasoning in real-world human-robot navigation. State-of-the-art VLMs were compared with human and rule-based baselines. The best VLM showed encouraging agreement with human answers, but still trailed simpler rule-based methods and human consensus. The authors frame that gap as a barrier to safe, socially compliant robot navigation. ArXiv · AI/CL/LG's note
The paper introduces SocialNav-SUB, a VQA benchmark for spatial, spatiotemporal, and social reasoning in real-world human-robot navigation. State-of-the-art VLMs were compared with human and rule-based baselines. The best VLM showed encouraging agreement with human answers, but still trailed simpler rule-based methods and human consensus. The authors frame that gap as a barrier to safe, socially compliant robot navigation. ArXiv · AI/CL/LG's note
score 4