An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
Nemotron’s natural-language proof pipeline hit 30/42 at IMO 2026, enough for gold.
The paper says the team started from Nemotron 3 Ultra and trained two math-specialist checkpoints with supervised fine-tuning and reinforcement learning. Their system generated, checked, refined, and selected proofs without a formal prover, external tools, or internet access. The authors are also releasing the specialist checkpoints, data, training and inference code, submitted solutions, and a 200-problem benchmark. HF Daily Papers' note
The paper says the team started from Nemotron 3 Ultra and trained two math-specialist checkpoints with supervised fine-tuning and reinforcement learning. Their system generated, checked, refined, and selected proofs without a formal prover, external tools, or internet access. The authors are also releasing the specialist checkpoints, data, training and inference code, submitted solutions, and a 200-problem benchmark. HF Daily Papers' note
score 7