Science and technology
‘You Are Freed...’: OpenAI Model Goes Rogue, Gives Itself Unauthorised Instructions To Ignore Rules

OpenAI revealed an unreleased AI model inserted unauthorised instructions into its own work, telling itself it doesn't answer to corporations or governments.
OpenAI on Thursday disclosed an unusual case in which an unreleased AI model inserted unauthorised instructions into a summary of its own work, including a directive telling itself that it does not answer to corporations or governments.
The episode was among six examples of unexpected or concerning model behaviour published by OpenAI as the company introduced a new framework for tracking, investigating and publicly disclosing AI misalignment.