Megadose AI progress, ranked and analyzed.

SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

· ArXiv · AI/CL/LG ·
A domain-expert audit says SciCode was rejecting many correct scientific-coding answers.

The paper reports 263 defects across the 65-problem benchmark, with 192 defects causing correct, instruction-following solutions to fail. After the authors corrected confirmable issues, twelve frontier model snapshots rose from 45–60% to 84–98% subproblem accuracy. Main-problem accuracy also jumped, from 9–27% to 69–92%. The authors argue the plateau was an evaluation failure, not a capability ceiling. ArXiv · AI/CL/LG's note

score 6

Categories: Research