PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
PhysBrain 1.5 trains a vision-language model to read scenes, plan robot motion, and predict what comes next in one framework.
The paper encodes language, end-effector movement, and dense visual targets as token sequences for autoregressive training. Its pre-training uses human interaction videos, then supervised fine-tuning mixes human demonstrations, robot trajectories, and simulation. The 8B model averages 72.5 across 28 embodied understanding benchmarks, which the authors call a new open-source state of the art. The report also shows qualitative examples of trajectory generation and future-scene prediction in RGB, depth, and robot-mask outputs. HF Daily Papers' note
The paper encodes language, end-effector movement, and dense visual targets as token sequences for autoregressive training. Its pre-training uses human interaction videos, then supervised fine-tuning mixes human demonstrations, robot trajectories, and simulation. The 8B model averages 72.5 across 28 embodied understanding benchmarks, which the authors call a new open-source state of the art. The report also shows qualitative examples of trajectory generation and future-scene prediction in RGB, depth, and robot-mask outputs. HF Daily Papers' note
score 6