Megadose AI progress, ranked daily.

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

· HF Daily Papers ·
OraRL treats annotations as oracle rollouts, making video MLLM post-training faster while improving reported benchmark scores.

The paper says standard RL post-training wastes compute on weak on-policy rollouts, especially with chain-of-thought generation. OraRL adds the annotation itself into the rollout group and uses a decoupled advantage estimator to avoid “advantage inversion.” The authors report 2.2x SFT step time, versus 4.9x for GRPO with CoT, and 130 ms decoding for Video-ORA-9B without CoT. Reported gains include higher temporal mIoU, tracking AO, segmentation, and VSI-Bench scores. HF Daily Papers' note

score 5

Categories: Research