How good are frontier models at physics?
Expert re-grading pushed leading physics benchmark scores sharply higher, arguing the tests were often wrong before the models were.
The paper audits six physics benchmarks with faculty and graduate researchers, focusing on text-only problems with verifiable answers. Many responses first marked incorrect were attributed to grader errors, bad reference solutions, or ambiguous questions. After correction, GPT-5.6-Sol rose from 47.3% to 78.7% mean@4 on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark. The authors say current closed-ended physics evaluations may be nearing saturation and need harder, expert-validated tests. HN · ArXiv's note
The paper audits six physics benchmarks with faculty and graduate researchers, focusing on text-only problems with verifiable answers. Many responses first marked incorrect were attributed to grader errors, bad reference solutions, or ambiguous questions. After correction, GPT-5.6-Sol rose from 47.3% to 78.7% mean@4 on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark. The authors say current closed-ended physics evaluations may be nearing saturation and need harder, expert-validated tests. HN · ArXiv's note
score 5
Discussions
- hn · 77 points · 36 comments
- hn · 80 points · 40 comments
- hn · 80 points · 41 comments
- hn · 81 points · 42 comments
- hn · 83 points · 42 comments
- hn · 83 points · 42 comments
- hn · 85 points · 43 comments
- hn · 86 points · 43 comments
- hn · 89 points · 43 comments
- hn · 89 points · 43 comments
- hn · 90 points · 44 comments
- hn · 90 points · 46 comments
- hn · 90 points · 46 comments
- hn · 90 points · 46 comments
- hn · 91 points · 46 comments
- hn · 91 points · 46 comments
- hn · 92 points · 47 comments
- hn · 92 points · 48 comments
- hn · 92 points · 48 comments
- hn · 94 points · 48 comments
- hn · 94 points · 48 comments