Patch Policy: Efficient Embodied Control via Dense Visual Representations
Patch Policy reports stronger robot control by using dense ViT patch tokens without carrying a full VLM.
The paper says its block-causal attention mask lets transformer policies attend to many visual patch tokens per observation while preserving temporal causality. Across four simulated and three real-world environment suites, it reports a 40% relative gain over policies using global-pooled representations. It also claims an 18% improvement over fine-tuned OpenVLA-OFT while using about 0.7% of the parameters. ArXiv · AI/CL/LG's note
The paper says its block-causal attention mask lets transformer policies attend to many visual patch tokens per observation while preserving temporal causality. Across four simulated and three real-world environment suites, it reports a 40% relative gain over policies using global-pooled representations. It also claims an 18% improvement over fine-tuned OpenVLA-OFT while using about 0.7% of the parameters. ArXiv · AI/CL/LG's note
score 5