PermVLA: Factorization Order as a Regularizer for VLA Learning
PermVLA treats action order itself as a training regularizer, then still runs standard left-to-right control at deployment.
The paper introduces causally anchored permutation, which samples different action reveal orders while keeping a tunable chronological prefix. That turns one expert action chunk into multiple subset-conditioned prediction tasks without adding demonstrations. In controlled tests, the authors report consistent gains over standard left-to-right training on LIBERO and LIBERO-Plus, with the advantage also appearing on CALVIN cross-dataset evaluation. They also report a diagnostic showing CAP-trained policies better agree across reveal orders. ArXiv · AI/CL/LG's note
The paper introduces causally anchored permutation, which samples different action reveal orders while keeping a tunable chronological prefix. That turns one expert action chunk into multiple subset-conditioned prediction tasks without adding demonstrations. In controlled tests, the authors report consistent gains over standard left-to-right training on LIBERO and LIBERO-Plus, with the advantage also appearing on CALVIN cross-dataset evaluation. They also report a diagnostic showing CAP-trained policies better agree across reveal orders. ArXiv · AI/CL/LG's note
score 4