Megadose AI progress, ranked and analyzed.

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

· HF Daily Papers ·
BridgeVLA++ adds spatio-temporal memory to a 3D robot manipulation framework built on pre-trained vision-language models.

The paper says the system keeps persistent spatial context and temporal interaction history, letting it reason over past observations. It is presented as retaining BridgeVLA’s data efficiency and generalization while improving memory-dependent manipulation. The authors report state-of-the-art results on two memory-focused benchmarks, plus effective bimanual manipulation and validation on another real-world robot platform. HF Daily Papers' note

score 5

Categories: Research