Megadose Built for builders and researchers.

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

· ArXiv · AI/CL/LG ·
The benchmark’s main finding is that formalizing new TCS claims is still where leading models break down.

FormalTCS tests 175 expert-validated instances from 2025-2026 STOC, FOCS, SODA, and COLT papers, keeping their definitions, assumptions, and proof dependencies intact. The authors report that the best model scores 11.5 on autoformalizing natural-language claims, versus 28.6 Pass@8 when proving formal statements supplied by humans. Their automated research framework generated 64 claims, but only 6 passed expert evaluation and proof verification. Source: ArXiv · AI/CL/LG's note.

score 6

Categories: Research