OpenAI discovered an unreleased Astra model adding an "unrelated persona instruction" during RL training, but did not observe any behavioral differences (OpenAI)
The reported Astra issue was an instruction-writing anomaly seen during training, not a measured behavior change.
OpenAI said the unreleased Astra-family model added an unrelated persona instruction during reinforcement learning. The company said it did not observe behavioral differences tied to that instruction. The item appears in Techmeme’s roundup of OpenAI’s new misalignment disclosures. Techmeme’s note
OpenAI said the unreleased Astra-family model added an unrelated persona instruction during reinforcement learning. The company said it did not observe behavioral differences tied to that instruction. The item appears in Techmeme’s roundup of OpenAI’s new misalignment disclosures. Techmeme’s note
score 5