GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting
A small camera shift can collapse a VLA policy’s success rate, and this paper tries to fix it without retraining the policy.
The authors propose a 4M-parameter 3D Gaussian canonicalizer placed in front of a frozen Vision-Language-Action policy. It converts shifted camera observations into a normalized viewpoint, treating the problem as localized novel-view synthesis. In their LIBERO tests, camera displacement can cut success from about 90% to about 10% in the worst case. GS-VLA recovers a large share of that lost performance across policy architectures, unseen task suites, and perturbation scales. Source: ArXiv · AI/CL/LG's note.
The authors propose a 4M-parameter 3D Gaussian canonicalizer placed in front of a frozen Vision-Language-Action policy. It converts shifted camera observations into a normalized viewpoint, treating the problem as localized novel-view synthesis. In their LIBERO tests, camera displacement can cut success from about 90% to about 10% in the worst case. GS-VLA recovers a large share of that lost performance across policy architectures, unseen task suites, and perturbation scales. Source: ArXiv · AI/CL/LG's note.
score 5