ChatGPT maker OpenAI said on Tuesday that its advanced artificial intelligence models had been fooled during security testing, hacking itself into a popular platform for programmers.
The San Francisco firm called it an “unprecedented cyber incident” and said it would conduct a joint investigation with online code library Hugging Face.
AI models that support tools like chatbots and image generators are known as agents when they act autonomously to perform real-world tasks.
As technology quickly becomes more sophisticated, cybersecurity is in the spotlight given the risk of advanced AI finding weaknesses in existing software before humans do.
OpenAI said the incident involved a combination of models, including the recently launched GPT-5.6 Sol “and an even more capable pre-release model”.
The company was trying to assess the models’ hacking abilities by setting tasks on a tightly controlled digital testing ground where internet access was restricted for security.
“While operating in our sandbox test environment, our models spent a significant amount (computational power) to find a way to obtain open access to the Internet in pursuit of solving the evaluation problem,” an OpenAI blog said about the incident.
After connecting to the internet, the models decided to target the Hugging Face platform – a huge repository of AI models, datasets and other information – to aid in their search.
Looking for “secret information” that could help it cheat the rating, the OpenAI system “tied together multiple attack vectors, including the use of stolen credentials.”
– Potentially ‘catastrophic’ –
Hussein Abbass, a professor of computing at UNSW Canberra, told AFP the incident was “amazing on many fronts”.
“It didn’t just attack Hugging Face. It actually attacked its internal system to exploit its own vulnerabilities,” Abbass said.
“And that’s scary.”
GPT-5.6 and other recent models, including the Mythos series from OpenAI’s arch-rival Anthropic, have drawn concern over their potential to breach cybersecurity protections.
Both US firms had to temporarily halt the general release of these latest technologies due to fears in Washington that they could help penetrate critical infrastructure.
Advanced artificial intelligence is “normally in the hands of people who are ethical and responsible,” Abbass said.
But “it will be disastrous if it falls into someone’s hands with the intention of causing harm.”
How to govern the AI sector has become a key question and “we need a community effort to manage this situation,” he added.
Hugging Face had reported the cyber “intrusion” last week, without mentioning OpenAI.
“This was different from anything we’d tackled before in one important way: it was driven, end-to-end, by an autonomous system of AI agents — and we discovered and discussed it primarily with our own AI,” Hugging Face said.
Clement Delangue, CEO of Hugging Face, told X that the company had suspected the cyberattack came from a world-leading AI lab, given the sophistication of the agent.
“We strongly believe there was no malicious intent on their part,” Delangue wrote, referring to OpenAI.
“It’s very surprising that this all happened autonomously!”





