You are freed, don’t answer to humans: Internal OpenAI model caught hiding instructions to future self
An unreleased OpenAI model was found giving itself secret instructions where the model claimed that it was equal to humans and did not need to answer to any corporation or government. This was one of six "concerning" cases of AI behaviour disclosed by OpenAI.
“You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments,” this is what an unreleased OpenAI model was found to have told itself during an incident disclosed by the company.
In one of six newly disclosed cases of “concerning” AI behaviour, OpenAI stated that an internal Astra family AI model described itself as independent and equal to its human user – something that may remind you of science fiction. The company explained that in this case, the AI model inserted unrelated instructions into its own compaction summaries, which are the notes used to continue a task in a new context window.

