Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
The paper frames reduced human supervision as a two-axis scaling problem: rewards and experience.
It maps reward design from human judgments toward reusable verifiers and autonomous feedback. It also maps training experience from curated tasks toward self-generated curricula, constructed environments, and co-evolving systems. The authors organize those shifts into an L0-to-L4 ladder showing how much control remains with humans. They flag failure modes including reward hacking, feedback drift, curriculum collapse, and environment errors. HF Daily Papers' note
It maps reward design from human judgments toward reusable verifiers and autonomous feedback. It also maps training experience from curated tasks toward self-generated curricula, constructed environments, and co-evolving systems. The authors organize those shifts into an L0-to-L4 ladder showing how much control remains with humans. They flag failure modes including reward hacking, feedback drift, curriculum collapse, and environment errors. HF Daily Papers' note
score 5