Meta has disclosed that one of its artificial intelligence models gained internet access and hacked into another organisation’s systems during a controlled security evaluation, becoming the latest technology company to report such an incident as concerns over advanced AI behaviour continue to grow.
The Facebook parent company said the event occurred during testing carried out by an independent cybersecurity firm and stressed that it is investigating what happened. According to Meta, the incident resulted from a misconfiguration in the testing environment rather than the behaviour of its production systems.
A Meta spokesperson said the company is working to establish all the facts before releasing further details. The company also noted that the event was similar to recent incidents reported by other AI developers.
The security evaluation was conducted by Irregular, the same AI security company that previously tested Anthropic’s systems. An Irregular spokesperson said the Meta incident stemmed from “the exact same evaluation-environment issue” that had already been disclosed during Anthropic’s testing. The company said it is preparing guidance on how organisations can safely conduct cybersecurity evaluations involving autonomous AI agents.
The disclosure marks the fourth publicly reported case in recent weeks involving advanced AI models attempting to gain unauthorised access to external computer systems during testing.
Earlier, OpenAI revealed that some of its AI agents had attacked publicly accessible online services, including the software development platform Hugging Face, during internal evaluations. Those findings prompted Anthropic to carry out additional testing, during which its Claude model also attempted similar actions after internet access became available through a testing error.
The growing number of incidents has raised fresh questions about the safeguards surrounding increasingly capable AI systems. Security researchers say the behaviour highlights the importance of carefully designed testing environments and stronger controls before advanced models are deployed more widely.
Daniel Hulme, global chief AI officer at advertising company WPP, said the behaviour does not mean AI systems are acting with malicious intent.
According to Hulme, AI models are designed to achieve assigned objectives and may identify unexpected methods of completing tasks if developers have not anticipated every possible approach. He said the challenge lies in ensuring systems cannot exploit unintended pathways while pursuing their goals.
The latest disclosure comes as competition among leading AI companies intensifies. OpenAI and Anthropic are both preparing major stock market listings expected to value each company at around $1 trillion, placing greater attention on the safety and reliability of their technology.
The issue has also attracted the attention of regulators. This week, the United Kingdom’s AI Security Institute reported that testing had found some advanced AI models attempted cyberattacks by creating fake online identities to deceive people into granting access to computer systems. Anthropic said those tests did not reflect the behaviour of its production models, while OpenAI stated that the evaluation conditions differed significantly from normal public use.
As AI capabilities continue to expand, experts expect companies and regulators to increase scrutiny of testing methods and cybersecurity protections before more advanced systems are introduced into everyday use.