FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
The benchmark’s main finding is that formalizing new TCS claims is still where leading models break down.
FormalTCS tests 175 expert-validated instances from 2025-2026 STOC, FOCS, SODA, and COLT papers, keeping their definitions, assumptions, and proof dependencies intact. The authors report that the best model scores 11.5 on autoformalizing natural-language claims, versus 28.6 Pass@8 when proving formal statements supplied by humans. Their automated research framework generated 64 claims, but only 6 passed expert evaluation and proof verification. Source: ArXiv · AI/CL/LG's note.
FormalTCS tests 175 expert-validated instances from 2025-2026 STOC, FOCS, SODA, and COLT papers, keeping their definitions, assumptions, and proof dependencies intact. The authors report that the best model scores 11.5 on autoformalizing natural-language claims, versus 28.6 Pass@8 when proving formal statements supplied by humans. Their automated research framework generated 64 claims, but only 6 passed expert evaluation and proof verification. Source: ArXiv · AI/CL/LG's note.
score 6