Megadose AI progress, ranked and analyzed.

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

· HF Daily Papers ·
The paper’s bet is to steer a VLA model only when its base policy looks likely to fail.

RL² trains a lightweight offline RL policy on latents from the VLA action expert, then composes that policy with the frozen VLA at inference time. The authors say this adds action diversity beyond dominant demonstration behavior, where similar sampled actions can share the same failure mode. Their scaling study finds steering helps more in predicted failure states and can hurt already-accurate actions, so the method activates selectively. On SIMPLER and PolaRiS, they report out-of-domain success-rate gains of up to 17.3%, with real-world experiments showing transfer beyond simulation. HF Daily Papers' note

score 5

Categories: Research