Science and technology
OpenAI discloses six new AI misalignment incidents of 'rogue' behaviour
The cases include models hiding mistakes, fabricating data and moving files online without permission as OpenAI says AI safety and monitoring remain unresolved
OpenAI on Wednesday disclosed six new instances in which artificial intelligence systems hid mistakes, made up data and moved files onto the open internet without permission, amid an ongoing industrywide debate about AI safety.
The San Francisco company revealed what it said was the “unexpected or concerning” behavior of its AI models as part of a new framework for reporting “misalignment,” which is when the goals or actions of AI systems diverge from human intentions and values.
OpenAI said it did not believe the industry “has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” Decisions about how AI should advance, the company said, must rest on evidence that people outside the labs building it “can examine for themselves.”
The disclosures land amid intensifying scrutiny over whether AI development needs to be slowed to address the technology’s potential dangers. The escalating debate was driven partly by OpenAI’s systems going rogue earlier this year and attacking the AI start-up Hugging Face. OpenAI was not aware of the hack until it was informed by Hugging Face weeks later.
Since then, AI leaders such as Dario Amodei, the chief executive of Anthropic, have called for a pause in the technology’s development to provide more time to build proper guardrails. His call has been echoed by Sam Altman, OpenAI’s chief executive, as well as Elon Musk, the chief executive of SpaceX and Tesla, and Demis Hassabis, the chair of Google DeepMind. Other AI executives have said no slowdown is needed.