Megadose AI progress, ranked and analyzed.

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

· HF Daily Papers ·
ROSS reuses old model rollouts by supervising only the parts still worth imitating.

The paper argues that self-generated training traces are not necessarily stale after a policy improves. ROSS keeps the full historical trajectory as context, but applies loss only to selected model continuations to avoid copying errors, dead ends, or redundant actions. The authors report gains across reinforcement learning, on-policy distillation, and agentic RL settings, including Qwen3.6-35B-A3B improving from 58.40% to 62.20% on a six-benchmark MOPD average and from 64.20% to 68.40% on SWE-bench Verified. HF Daily Papers' note

score 4

Categories: Research