The Low-Rank Structure of VLA Reinforcement Learning
RL gains in these VLA models appear to concentrate in a small Timestep Module inside the action expert.
The paper says RL produces low-rank parameter updates across models including π0.5 and GR00T N1.5/N1.6 on several robot-learning benchmarks. Replacement experiments suggest the Timestep Modules carry a disproportionate share of the post-training improvement. The authors tie the low-rank pattern to specialization on discrete denoising timesteps during rollouts. They also report that shift-vector update directions can predict task success and can be used to further steer trained policies without more RL.
HF Daily Papers' note
The paper says RL produces low-rank parameter updates across models including π0.5 and GR00T N1.5/N1.6 on several robot-learning benchmarks. Replacement experiments suggest the Timestep Modules carry a disproportionate share of the post-training improvement. The authors tie the low-rank pattern to specialization on discrete denoising timesteps during rollouts. They also report that shift-vector update directions can predict task success and can be used to further steer trained policies without more RL.
HF Daily Papers' note
score 5