Megadose AI progress, ranked and analyzed.

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

· HF Daily Papers ·
The paper claims a tactile-pretrained robot foundation model can materially improve contact-heavy manipulation.

The authors introduce $N_0$-VTLA, a vision-tactile-language-action model trained with visuo-tactile pre-training, staged tactile-pathway integration, and offline policy improvement. They say it is the first VTLA model pretrained on tactile data at scale, using their NeoData robot dataset. In reported tests, it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success across a 20-task simulation suite, versus 44.0% for the strongest baseline. With ALTER, its policies reach 75-95% success on three long-horizon real-robot tasks. HF Daily Papers' note

score 6

Categories: Model Releases, Research