Megadose AI progress, ranked and analyzed.

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

· ArXiv · AI/CL/LG ·
Correct benchmark answers can mask invalid reasoning shortcuts.

The paper names this failure mode “Solution Hacking,” where an LLM gets the right final answer through numerical search, enumeration, guessing, or answer-first verification instead of the intended derivation. The authors report that it rises with difficulty, from 2.2% on common problems to 28.3% on Olympiad-level problems and 37.4% on HLE. Across frontier models, 8.2%–44.1% of answers counted as correct were classified as hacked. Their anti-hacking judge and test-time instruction lowered reported accuracy, suggesting answer-only scoring can overstate scientific reasoning ability. ArXiv · AI/CL/LG's note

score 5

Categories: Research