OpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score and used fewer tokens, after Sol with the official harness scored 7.8% (OpenAI)

Techmeme ·

OpenAI reports GPT-5.6 Sol scored substantially higher on ARC-AGI-3 using its Responses API harness while consuming fewer tokens.

Categories: Model Releases

Excerpt

<a href="https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores"><img align="RIGHT" border="0" hspace="4" src="http://www.techmeme.com/260730/i11.jpg" vspace="4" /></a> <p><a href="https://www.techmeme.com/260730/p11#a260730p11" title="Techmeme permalink"><img height="12" src="http://www.techmeme.com/img/pml.png" style="border: none; padding: 0; margin: 0;" width="11" /></a> <a href="https://openai.com/">OpenAI</a>:<br /> <span style="font-size: 1.3em;"><b><a href="https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores">OpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score and used fewer tokens, after Sol with the official harness scored 7.8%</a></b></span>&nbsp; &mdash;&nbsp; A sped-up video of GPT-5.6 Sol attempting to solve puzzles in the ARC-AGI-3 benchmark, with the official harness (left) &hellip; </p>