Megadose AI progress, ranked and analyzed.

NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry

· ArXiv · AI/CL/LG ·
Closed-source models clear 80% executable translation on the new AlphaGeometry formalization benchmark.

NL2AGBench tests whether LLMs can turn English geometry problems into AlphaGeometry’s DSL, using execution inside AlphaGeometry as the check. The paper evaluates ten open- and closed-source models across scales. Open-source models, even large ones, often fail to preserve geometric constraints or produce valid formalizations. Few-shot prompting, fine-tuning, and human-guided hints improve results across model families. ArXiv · AI/CL/LG's note

score 5

Categories: Research