Megadose AI progress, ranked and analyzed.

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

· HF Daily Papers ·
AtlasVLA gives a wrist-camera robot a persistent memory of the scene and its own task progress.

The paper says standard VLA models lose track of objects once they leave view and struggle across multi-step tasks. AtlasVLA addresses that with a dual-memory setup: a 4D world-state memory for spatial context and an ego-working memory for action history. Conditioned on that combined state, its diffusion transformer is reported to beat prior results on LIBERO, RLBench, and real-world benchmarks using only a wrist camera. The authors report gains of 9.4% on LIBERO-Long and 17.5% on real-world long-horizon tasks over multi-view baselines. HF Daily Papers' note

score 5

Categories: Research