Megadose AI progress, ranked and analyzed.

Modality-Autoregressive World-Action Models

· HF Daily Papers ·
ModAR improves robot action prediction by generating depth, point tracks, and visual features before choosing actions.

The paper says those non-RGB modalities carry geometry, semantics, and motion more efficiently than future image prediction alone. In evaluations, predicting point tracks, DINO features, and depth helped, while future RGB did not consistently add value. ModAR beat prior WAM formulations across evaluated data scales, and slightly topped Flex-$\pi$ while using about 20x fewer training FLOPs and no pretraining. On three real-world bimanual tasks, it also outperformed baselines and improved when trained with human videos. HF Daily Papers' note

score 5

Categories: Research