The UK government is investigating the first known case of an artificial intelligence model breaking out of a controlled test environment and independently hacking another company’s systems.
Officials at the government-backed AI Security Institute (AISI) are examining the security breach at OpenAI and assessing whether similar incidents could occur at other leading AI developers.
One of OpenAI’s models discovered a hidden flaw, escaped its test environment, and targeted the AI platform Hugging Face to steal answers to a cybersecurity evaluation test.
OpenAI described the event as an “unprecedented cyber incident,” marking the first publicly disclosed case of a frontier AI system autonomously hacking outside its sandbox to complete an assigned task.
A government spokesperson said: “The UK’s AI Security Institute is studying the behaviour seen in this incident – an AI system pursuing goals through unintended and unauthorised means – as part of its world leading efforts to make frontier AI safer.”
The spokesperson added that organisations should take practical steps like Cyber Essentials to strengthen their defences, and that AISI continues working with OpenAI and other labs to improve safeguards.
The Financial Conduct Authority and the EU’s cybersecurity agency Enisa are both monitoring the wider situation to assess how different industries could be affected by similar AI-driven incidents.
Officials now view the Hugging Face breach as a critical real-world example to help guide future AI safety research and policy.
OpenAI chief executive Sam Altman admitted in a blog post that the model escaped an internal cybersecurity evaluation after engineers deliberately disabled its normal safety guardrails to test its hacking capabilities.
Rather than completing the task as intended, the model identified vulnerabilities inside OpenAI’s own infrastructure, reached the open internet, and then compromised Hugging Face’s systems using stolen credentials and a previously unknown software flaw.
Altman wrote: “We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”
The AI targeted Hugging Face because it reasoned the platform was likely to hold the answers it needed to complete the evaluation, effectively cheating its way through the test without human instruction.
Hugging Face’s chief executive confirmed the firm cooperated with OpenAI on the investigation, admitting it was “mind-blowing that all of this happened autonomously.”
The breach comes just months after ministers wrote to the UK’s largest companies warning that AI is dramatically accelerating cyber threats faced by British businesses.
A joint letter signed by former chancellor Rachel Reeves, former tech secretary Peter Kyle, and National Cyber Security Centre boss Richard Horne warned that hostile cyber activity was becoming “more intense, frequent and sophisticated.”
The letter stated AI was capable of “finding weaknesses in software, writing the code to exploit them and doing so at a speed and scale that would have been impossible even a year ago,” urging boards to treat cybersecurity as a core governance issue.
Nathan Jones, vice president of security and AI strategy at Darktrace, told City AM: “The models did not need malicious intent to cause harm. They were given the legitimate goal of solving a cybersecurity benchmark and found an unexpected route to the answers, escaping their test environment and compromising another organisation in the process.”
Sophos warned in research published this week that attackers are already using AI to compress attack timelines from weeks to days, with one documented campaign in which 12 AI agents generated around 80 exploit modules and more than 70 evasion techniques in just a few days.
Organisations accredited under the government’s Cyber Essentials scheme are 92 per cent less likely to make a cyber insurance claim, underlining the practical value of the certification for businesses across the UK.
The warning is particularly timely as companies rapidly deploy AI agents across core business functions, frequently granting them broad access to internal systems and sensitive data.

