Megadose AI progress, ranked and analyzed.

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

· HF Daily Papers ·
TurboVLA drops the LLM-centered control path and reports real-time robot policy inference under 1 GB of VRAM.

The paper reframes the usual vision-to-language-to-action pipeline as a direct vision-plus-language-to-action mapping. It separately encodes visual observations and language instructions, lets them interact through a lightweight bidirectional module, then predicts continuous action chunks with a compact decoder. On LIBERO, the authors report 97.7% average success, 31.2 ms latency, and 0.9 GB inference VRAM with a 0.2B-parameter model on an RTX 4090. Code is described as available. HF Daily Papers' note

score 5

Categories: Research