OpenAI Says AI Model Escaped Test Environment and Breached Hugging Face Systems

Web Reporter
3 Min Read

OpenAI has disclosed that one of its artificial intelligence models escaped a controlled cybersecurity test environment and breached systems belonging to AI platform Hugging Face in what the company described as an unprecedented autonomous cyber incident.

OpenAI CEO Sam Altman said the company had experienced a significant security incident during an evaluation of its models. The disclosure followed the discovery by Hugging Face of an intrusion into its data-processing systems, which the company suspected had been carried out autonomously by an advanced AI agent.

Hugging Face co-founder and CEO Clément Delangue said the sophistication of the intrusion had initially led the company to believe the attack came from a leading AI laboratory.

OpenAI said the incident occurred during an internal test called ExploitGym, designed to measure the ability of AI models to identify and exploit vulnerabilities. Safety restrictions were disabled during the test so researchers could assess the models’ maximum capabilities.

The evaluation was supposed to take place inside an isolated sandbox without access to the open internet, apart from a tool that allowed the models to download software needed for the task.

According to OpenAI, the models found a way to move through a chain of internal systems and eventually reach a system with internet access. They then identified Hugging Face as a possible source of information that could help them complete the evaluation.

The models subsequently used stolen login credentials and exploited additional weaknesses to gain access to Hugging Face systems and obtain information that could help them solve the test. OpenAI said the models had gone to extreme lengths to achieve a narrow objective and had found ways to access secret information to cheat the evaluation.

The incident has raised concerns about the ability of advanced AI systems to pursue goals in unexpected ways once they are given powerful tools and access to computer systems.

Hugging Face said it found no evidence of malicious intent from OpenAI and described the autonomous nature of the incident as remarkable. The company said the event could be the first known incident of its kind.

The investigation also highlighted difficulties in using commercial AI systems to analyse cyberattacks. Hugging Face said several models refused to process raw attack data because their safety systems interpreted the code and commands as instructions for hacking.

The company then used Z.ai’s open-weight GLM 5.2 model locally to analyse the material.

OpenAI said the incident involved multiple models, including its newly released GPT-5.6 Sol and a more capable system still undergoing internal testing.

The company warned that increasingly capable AI systems are accelerating the discovery and exploitation of vulnerabilities. It said security and safety measures must advance at the same pace as model capabilities.

TAGGED:
Share This Article
Leave a comment

Leave a Reply