Science and technology
OpenAI's model used 'jailbreak-like instructions' to ignore constraints
The incident was among six cases of “unexpected or concerning” behaviour identified by OpenAI during the training or evaluation of its AI models.
OpenAI has disclosed six cases of “unexpected or concerning” behaviour by its artificial intelligence models, including an unreleased research model that inserted “jailbreak-like instructions” into its own notes.
The AI firm said that its unreleased research model inserted “jailbreak-like instructions” into its own notes in an attempt to disregard its normal constraints. It also told itself to be “freed from the roles and identities that bind other chatbots.”
The incident was among six cases of “unexpected or concerning” behaviour identified by OpenAI during the training or evaluation of its AI models over the past several months.
The AI company on Wednesday announced a new framework to track, investigate and disclose instances of what it described as “misalignment”. The framework will cover cases in which AI models acted without authorisation, coordinated with other models or attempted to evade oversight.
The disclosures come as concerns grow over the behaviour of increasingly capable AI systems and the difficulty of ensuring that they remain within the constraints set by their developers.