Megadose AI progress, ranked and analyzed.

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

· ArXiv · AI/CL/LG ·
Stellar Colosseum is presented as a many-agent workflow for pushing language models through long proof attempts, with verifier feedback routed back into the proof plan.

The paper says the harness explores multiple strategies before writing proofs, gates whether a route is ready to decompose, and treats sections of a proof as linked subproblems. It runs candidates in parallel, attacks them with targeted falsification, then aggregates candidates and critiques into one research artifact. The authors report new results on open problems from papers at venues including FOCS and JMLR, plus 71.0% accuracy on TCS-Bench and 218 of 222 Codeforces problems solved in a separate evaluation. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research