DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination
DeCAL reports a 71% average success rate on contact-rich dexterous manipulation tasks.
The paper presents a vision-language-action model built to handle occlusions and contact dynamics by using tactile signals more selectively. Its Mixture-of-Transformers setup separates experts for understanding, imagination, and action generation while allowing information to move between them. The authors add contact-aware visuo-tactile fusion and latent co-imagination to model visual and tactile dynamics together. They say the system reaches state-of-the-art results across all tested tasks, with an 83.4% progress success rate and generalization to unseen scenarios. ArXiv · AI/CL/LG's note
The paper presents a vision-language-action model built to handle occlusions and contact dynamics by using tactile signals more selectively. Its Mixture-of-Transformers setup separates experts for understanding, imagination, and action generation while allowing information to move between them. The authors add contact-aware visuo-tactile fusion and latent co-imagination to model visual and tactile dynamics together. They say the system reaches state-of-the-art results across all tested tasks, with an 83.4% progress success rate and generalization to unseen scenarios. ArXiv · AI/CL/LG's note
score 5