The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
A new benchmark tests whether AI systems can build better mathematical objects, not just answer fixed problems.
Muhan Zhang’s paper introduces the “Endless Exam,” built from fourteen parameterized construction families. Submissions are automatically checked and scored against published frontiers or baselines, with no ceiling on improvements. In tests across eight models and 69 instances, the scores separated model performance, but none beat a published frontier. The release includes generators, verifiers, references, model outputs, and analysis for continued benchmarking. ArXiv · AI/CL/LG's note
Muhan Zhang’s paper introduces the “Endless Exam,” built from fourteen parameterized construction families. Submissions are automatically checked and scored against published frontiers or baselines, with no ceiling on improvements. In tests across eight models and 69 instances, the scores separated model performance, but none beat a published frontier. The release includes generators, verifiers, references, model outputs, and analysis for continued benchmarking. ArXiv · AI/CL/LG's note
score 5