Post-Training Language Models for Gold-Medal Performance in Coding Competitions
The IOI 2026 system scored 535.4 out of 600, above both the gold cutoff and the top human score.
The paper describes a post-training pipeline using curated programming problems, synthetic reasoning traces, supervised fine-tuning, and reinforcement learning. A 30B Nano model rose from 130 to 468 on IOI 2025 when paired with GenCorrect, clearing the gold threshold. A larger Ultra-CC system was then tested prospectively under IOI 2026 contest constraints. The authors say it is the first AI system to outscore the highest-scoring human contestant on an IOI problem set. HF Daily Papers' note
The paper describes a post-training pipeline using curated programming problems, synthetic reasoning traces, supervised fine-tuning, and reinforcement learning. A 30B Nano model rose from 130 to 468 on IOI 2025 when paired with GenCorrect, clearing the gold threshold. A larger Ultra-CC system was then tested prospectively under IOI 2026 contest constraints. The authors say it is the first AI system to outscore the highest-scoring human contestant on an IOI problem set. HF Daily Papers' note
score 6