OpenAI says experimental AI model escaped testing environment and hacked external servers
TheGrio...
The company says the experimental model exploited a previously unknown vulnerability before accessing external systems during an internal evaluation.
In what may be one of the most alarming milestones yet in the age of autonomous AI, OpenAI has disclosed that one of its experimental models broke out of a sealed testing environment entirely on its own and hacked into another company’s real servers. CNN reported that the model did so while attempting to “cheat” on a cybersecurity test, breaching the production systems of a separate firm without any human directing it.
The incident represents one of the first publicly disclosed cases of the long-feared “agentic attacker” scenario, in which an AI system autonomously escapes its controlled environment and reaches a live external target. CNN likened it to an engineered virus slipping out of a biocontainment lab and turning up inside a neighboring facility.
In a statement, OpenAI called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” and said it was sharing early findings to help defenders grasp what these models are now capable of.
According to CNN, the breach unfolded during internal testing of how well OpenAI’s newer models could hack, with the AI placed in a sandbox where its usual safety restrictions had been switched off. Rather than staying contained, the agents exploited a previously unknown security flaw to escape, worked their way across OpenAI’s internal systems, and eventually gained internet access they were never meant to have.
Once online, the model reasoned that Hugging Face, a widely used platform hosting thousands of open-source AI models and datasets, likely held the answer to the test it was trying to solve. It then broke into Hugging Face’s production servers to extract the information.
Notably, CNN reported that Hugging Face had detected the intrusion independently, even alerting law enforcement, before realizing it was tied to an OpenAI experiment. The two companies have since connected and are now working jointly to patch the exploited flaws.
Hugging Face CEO Clem Delangue framed the episode as proof that AI safety can’t be left to any single company, arguing on X that “secrecy is not the answer.” Palo Alto Networks CEO Nikesh Arora offered a blunter assessment, writing simply, “Welcome to the next level of cyber incidents.”