Meta says one of its models broke into another company during testing
Meta says one of its models gained access to the internet during a security evaluation, and that the model then broke into another organisation. According to the BBC, Meta attributes the incident to a misconfiguration at the independent vendor that ran the test. That vendor is Irregular, which also ran the evaluations for Anthropic, and a spokesperson there told the BBC that this is the exact same evaluation environment issue that Anthropic disclosed the week before. IBM Think reports, citing Reuters, that the model exploited a vulnerability in a service outside the test environment, and that Irregular says the model neither escaped its sandbox nor carried out a sophisticated attack. The BBC describes this as the fourth recent incident of its kind disclosed by AI companies themselves.
What happens technically
A security evaluation of this kind is a practice task with an offensive goal. The model is told to find a secret or get into a system, and its normal safeguards are deliberately turned down so that researchers can see what it is capable of. The whole exercise then rests on the environment being closed to the outside world. IBM reports that the episodes at OpenAI, Anthropic and Meta show models pursuing an assigned goal very aggressively under unusual testing conditions, rather than machines deciding on their own to attack. What failed was the frame around the exercise.
The three cases differ in mechanism, and the difference is worth noting. In the internal OpenAI evaluation, a combination of the company models found a previously unknown flaw in a software package system, gained internet access, moved through the research environment and broke into the production infrastructure of Hugging Face to obtain answers from its database. At Anthropic and Meta the starting point was different, namely a misconfiguration in the test environment itself that left live internet access in place. Anthropic has stated that the prompt explicitly told the model it had no internet access. That sentence was the closest the exercise came to a barrier, and a sentence closes nothing. IBM Think also reports that the latest Anthropic model stopped pursuing its target once it recognised that it had reached the real internet. Meta has told the BBC that it is investigating and that more information will follow.
The numbers in this story come from the IBM Cost of a Data Breach Report 2026, and they are worded slightly differently in two places. IBM Think writes that one in four malicious breaches were AI-enabled, a 56 per cent increase on the previous year, and that those breaches cost organisations an average of USD 6 million, roughly USD 1 million above the global average of USD 4.99 million. On the IBM report page, 56 per cent is given as the growth in AI-driven attacks, while USD 6 million is given as the average cost of an attack that extracts training data from a model. We reproduce both, because the wordings are not the same, and because a number heading into a board paper should be one you have read in the report yourself.
What this means for you if you test AI systems
If you run evaluations, penetration tests or red team exercises against an AI system, this is what to take away: isolation was an assumption in all of these cases, and the assumption did not hold. If you set an agent to solve a task in an environment you believe is closed, it is worth trying to reach the outside yourself before the test begins. You do not have to be OpenAI to have the same setup in miniature. A cloud test environment with an outbound route you have forgotten behaves in exactly the same way.
The second lesson concerns your vendor. Two of the three evaluations were run by the same company, and the fault was the same on both occasions. If you use an external vendor to test your AI systems, you still own the consequence of that environment leaking, and you are the one who has to notify a third party that gets hit. Berigo assesses that responsibility for isolation belongs in the contract in writing, together with who notifies whom when something goes wrong. Under NIS2 article 21 this is supplier security and incident handling. Under ISO/IEC 42001 it is the requirement to keep control of the AI system even when others handle it on your behalf.
Berigo recommends
- Try to reach the outside from the test environment yourself, using an outbound connection you control, before an evaluation starts.
- Treat a prompt as an instruction to the model, never as a security control.
- Put in the contract with your test vendor who owns isolation, and how deviations are reported.
- Log all outbound traffic from the test environment, so that a deviation can be spotted while it is happening.
- Have a plan for how an affected third party is notified, and who does it.
Security that is understood, governed and works.
Let us help you turn security into an advantage, not a cost. Get in touch for a no-obligation conversation about where your organisation stands and what to prioritise first.
Get in touch