Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Peer-rewarded cohorts improved reasoning without ground-truth labels.
Co-RL trains multiple decoupled models with rewards derived from one another’s outputs, rather than verified human labels. The paper says diversity across model families, sizes, and rephrased samples reduces correlated errors and helps avoid collapse. Reported gains are 3.0-8.6% across seven text-only benchmarks and 2.3-7.2% across four multimodal benchmarks. HF Daily Papers' note
Co-RL trains multiple decoupled models with rewards derived from one another’s outputs, rather than verified human labels. The paper says diversity across model families, sizes, and rephrased samples reduces correlated errors and helps avoid collapse. Reported gains are 3.0-8.6% across seven text-only benchmarks and 2.3-7.2% across four multimodal benchmarks. HF Daily Papers' note
score 5