EnigmaForge: The Question Is Hidden in the Story
EnigmaForge tests whether models can infer the task before solving it.
The benchmark gives models generated document bundles containing a hidden, uniquely solvable logic puzzle. Its generator proves uniqueness with a SAT solver and checks that each clue matters. Across 25 frontier models and 17,400 scored records, performance on “intuition” varied far more than factual recovery. The abstract also notes that some models failed because content filters stopped them before the puzzle. ArXiv · AI/CL/LG's note
The benchmark gives models generated document bundles containing a hidden, uniquely solvable logic puzzle. Its generator proves uniqueness with a SAT solver and checks that each clue matters. Across 25 frontier models and 17,400 scored records, performance on “intuition” varied far more than factual recovery. The abstract also notes that some models failed because content filters stopped them before the puzzle. ArXiv · AI/CL/LG's note
score 6