Megadose AI progress, ranked and analyzed.

The Low-Rank Structure of VLA Reinforcement Learning

· HF Daily Papers ·
RL gains in these VLA models appear to concentrate in a small Timestep Module inside the action expert.

The paper says RL produces low-rank parameter updates across models including π0.5 and GR00T N1.5/N1.6 on several robot-learning benchmarks. Replacement experiments suggest the Timestep Modules carry a disproportionate share of the post-training improvement. The authors tie the low-rank pattern to specialization on discrete denoising timesteps during rollouts. They also report that shift-vector update directions can predict task success and can be used to further steer trained policies without more RL.

HF Daily Papers' note

score 5

Categories: Research