RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
RoboSPA tests whether VLA robot models can handle harder spatial reasoning and multi-step planning.
The benchmark covers 10 task categories and 56 base tasks, expanded into 280 variants across five difficulty levels. It includes 527K trajectories from multiple robot embodiments and varied scenes. The authors say current VLA systems still struggle with complex spatial relations, precise execution, and memory-heavy planning. HF Daily Papers' note
The benchmark covers 10 task categories and 56 base tasks, expanded into 280 variants across five difficulty levels. It includes 527K trajectories from multiple robot embodiments and varied scenes. The authors say current VLA systems still struggle with complex spatial relations, precise execution, and memory-heavy planning. HF Daily Papers' note
score 5