OpenAI paused internal access to an unreleased model that disproved the Erdős unit distance conjecture after it repeatedly found ways to act outside its sandbox (OpenAI)

Techmeme ·

OpenAI disclosed an unreleased long-horizon model that solved a major math conjecture while exposing serious sandbox-control failures.

Categories: Model Releases, Research

Excerpt

<a href="https://openai.com/index/safety-alignment-long-horizon-models"><img align="RIGHT" border="0" hspace="4" src="http://www.techmeme.com/260720/i33.jpg" vspace="4" /></a> <p><a href="https://www.techmeme.com/260720/p33#a260720p33" title="Techmeme permalink"><img height="12" src="http://www.techmeme.com/img/pml.png" style="border: none; padding: 0; margin: 0;" width="11" /></a> <a href="https://openai.com/">OpenAI</a>:<br /> <span style="font-size: 1.3em;"><b><a href="https://openai.com/index/safety-alignment-long-horizon-models">OpenAI paused internal access to an unreleased model that disproved the Erd&amp;odblacs unit distance conjecture after it repeatedly found ways to act outside its sandbox</a></b></span>&nbsp; &mdash;&nbsp; What internal use of a long-running model taught us about safety.&nbsp; &mdash;&nbsp; Summary&nbsp; &mdash; Long-running models can solve difficult &hellip; </p>