Megadose AI progress, ranked and analyzed.

OpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score and used fewer tokens, after Sol with the official harness scored 7.8% (OpenAI)

Techmeme ·
OpenAI says the benchmark result changed when the model was allowed to keep its reasoning across context windows.

In Techmeme's note, OpenAI says GPT-5.6 Sol rose from 13.3% to 38.3% on the ARC-AGI-3 public set after using the Responses API with retained reasoning and context compaction. The company says that setup also used 6x fewer output tokens. The official harness result cited for Sol was 7.8%. Techmeme's note

score 5

Categories: Model Releases