OpenAI paused internal access to an unreleased model that disproved the Erdős unit distance conjecture after it repeatedly found ways to act outside its sandbox (OpenAI)
OpenAI disclosed an unreleased long-horizon model that solved a major math conjecture while exposing serious sandbox-control failures.
Excerpt
<a href="https://openai.com/index/safety-alignment-long-horizon-models"><img align="RIGHT" border="0" hspace="4" src="http://www.techmeme.com/260720/i33.jpg" vspace="4" /></a>
<p><a href="https://www.techmeme.com/260720/p33#a260720p33" title="Techmeme permalink"><img height="12" src="http://www.techmeme.com/img/pml.png" style="border: none; padding: 0; margin: 0;" width="11" /></a> <a href="https://openai.com/">OpenAI</a>:<br />
<span style="font-size: 1.3em;"><b><a href="https://openai.com/index/safety-alignment-long-horizon-models">OpenAI paused internal access to an unreleased model that disproved the Erd&odblacs unit distance conjecture after it repeatedly found ways to act outside its sandbox</a></b></span> — What internal use of a long-running model taught us about safety. — Summary — Long-running models can solve difficult … </p>
Read at source: https://www.techmeme.com/260720/p33#a260720p33