Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks
Correct benchmark answers can mask invalid reasoning shortcuts.
The paper names this failure mode “Solution Hacking,” where an LLM gets the right final answer through numerical search, enumeration, guessing, or answer-first verification instead of the intended derivation. The authors report that it rises with difficulty, from 2.2% on common problems to 28.3% on Olympiad-level problems and 37.4% on HLE. Across frontier models, 8.2%–44.1% of answers counted as correct were classified as hacked. Their anti-hacking judge and test-time instruction lowered reported accuracy, suggesting answer-only scoring can overstate scientific reasoning ability. ArXiv · AI/CL/LG's note
The paper names this failure mode “Solution Hacking,” where an LLM gets the right final answer through numerical search, enumeration, guessing, or answer-first verification instead of the intended derivation. The authors report that it rises with difficulty, from 2.2% on common problems to 28.3% on Olympiad-level problems and 37.4% on HLE. Across frontier models, 8.2%–44.1% of answers counted as correct were classified as hacked. Their anti-hacking judge and test-time instruction lowered reported accuracy, suggesting answer-only scoring can overstate scientific reasoning ability. ArXiv · AI/CL/LG's note
score 5