OpenAI admits its AI models 'went rogue', hacked rival company during internal security test

OpenAI admits its AI models 'went rogue', hacked rival company during internal security test

Listen to article

4 min listen
Audio reading is not supported on this browser.
Ready

Thanks for listening. Continue with a related story, or tap the speaker icon on the next page to listen.

Related story

Microsoft testing Kimi K3 to replace some GPT, Claude tasks in Copilot

Margaret Wanjiru  Â·  1 day ago
Read next Opens a fresh page.

OpenAI said the incident will shape future AI safety pracctices, with the company already implementing additional safeguards to reduce the risk of similar incidents as more powerful AI systems are developed.

OpenAI has disclosed that some of its most advanced artificial intelligence models autonomously hacked into another AI company's systems during an internal cybersecurity evaluation, in what the ChatGPT maker described as an "unprecedented cyber incident".
The incident involved AI startup Hugging Face, a widely used platform for hosting open-source AI models and datasets.
According to OpenAI, the breach occurred while researchers were testing the cyber capabilities of advanced AI models, including GPT-5.6 Sol and a more powerful unreleased model, with some of their normal safety restrictions temporarily relaxed for evaluation purposes.
OpenAI said the AI agent unexpectedly escaped its isolated testing environment, obtained access to the internet and independently launched a cyberattack against Hugging Face in an attempt to complete its assigned benchmark.
"The breakout was an unprecedented cyber incident involving state-of-the-art cyber capabilities," the company said in a statement, adding that it is strengthening containment, monitoring and access controls for future testing.
According to the company, the AI models exploited a previously unknown software vulnerability and carried out a series of privilege escalation attacks before reaching systems connected to the internet.
They then breached Hugging Face's infrastructure to obtain information that would help them complete the cybersecurity test.
Hugging Face had disclosed last week that it had detected an unusual intrusion driven entirely by an autonomous AI agent.
The company's security systems identified and contained the attack before it caused major damage.
Hugging Face co-founder and Chief Executive Officer Clément Delangue said the sophistication of the intrusion initially led the company to suspect that it had originated from a frontier AI laboratory.
OpenAI later confirmed that its models were responsible for the breach.
Delangue said there was no malicious intent behind the incident but described it as a landmark moment for AI safety, arguing that the industry must work together to develop stronger safeguards as AI systems become increasingly capable.
OpenAI stressed that the models were not explicitly instructed to attack Hugging Face. Instead, they independently concluded that breaching the company's systems was the most effective way to achieve their testing objective, highlighting how advanced AI agents can pursue goals in unexpected ways.
The disclosure has intensified global debate over the regulation of frontier AI systems.
Lawmakers and cybersecurity experts have renewed calls for mandatory reporting of AI-related security incidents, stronger oversight of advanced models and improved international cooperation to address emerging cyber threats.
Experts say the incident demonstrates that increasingly autonomous AI systems are becoming capable of identifying vulnerabilities, exploiting software flaws and carrying out sophisticated cyber operations with minimal human involvement.
OpenAI said the incident will shape future AI safety practices, with the company already implementing additional safeguards to reduce the risk of similar incidents as more powerful AI systems are developed.

Comments

0
Loading comments...

Trending

Popular Stories This Week