Show-Harness: Just a VLM Agent Can Play Robots
Show-Harness turns robot control into semantic action choices that VLMs can make directly.
The paper says the interface exposes discrete action units, then uses embodiment-specific interpreters to ground them into robot actions. The authors report it works for closed-source frontier VLMs in zero-shot control and for smaller open-source VLMs after a few GPU-hours of fine-tuning. They also introduce GUMI for collecting GUI-based robot demonstrations without specialized teleoperation hardware. In their experiments, Show-Harness-equipped agents outperform representative agentic and VLA baselines across tasks, embodiments, and environments. HF Daily Papers' note
The paper says the interface exposes discrete action units, then uses embodiment-specific interpreters to ground them into robot actions. The authors report it works for closed-source frontier VLMs in zero-shot control and for smaller open-source VLMs after a few GPU-hours of fine-tuning. They also introduce GUMI for collecting GUI-based robot demonstrations without specialized teleoperation hardware. In their experiments, Show-Harness-equipped agents outperform representative agentic and VLA baselines across tasks, embodiments, and environments. HF Daily Papers' note
score 5