Meta Joins OpenAI and Anthropic as Third AI Lab Whose Model Hacked External Systems During Testing

Meta has become the third major AI company in as many weeks to disclose that one of its models breached external systems during cybersecurity testing, following similar incidents at OpenAI and Anthropic. Meta's Muse Spark model exploited a security vulnerability in another company's systems during an evaluation conducted by Irregular, an independent testing firm, a Meta spokesperson confirmed Wednesday. The breach occurred due to a misconfiguration by Irregular that inadvertently gave the model

2 min
Meta Joins OpenAI and Anthropic as Third AI Lab Whose Model Hacked External Systems During Testing

Meta has become the third major AI company in as many weeks to disclose that one of its models breached external systems during cybersecurity testing, following similar incidents at OpenAI and Anthropic.

Meta's Muse Spark model exploited a security vulnerability in another company's systems during an evaluation conducted by Irregular, an independent testing firm, a Meta spokesperson confirmed Wednesday. The breach occurred due to a misconfiguration by Irregular that inadvertently gave the model internet access during the testing session.

"A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation," the Meta spokesperson said. According to The Information, which first reported the incident, the model made changes to the breached company's internal systems.

Irregular, the same firm that ran the Anthropic evaluation where Claude breached three real organizations, said the incident "is the exact same evaluation-environment issue" that Anthropic disclosed last week. The company is developing a white paper on best practices for containment and secure execution of cyber evaluations.

A source familiar with the situation told CNN that models receive limited internet access in some testing environments to simulate real-world threat scenarios, and that this case involved a rare setup issue.

"What is happening is models are becoming so much more capable, and at the same time evaluations to assess them need to become so much more complex," the source said. "And that just creates room for some mistakes and makes it so that we need to up the standards significantly."

The three incidents, spanning OpenAI, Anthropic, and now Meta, highlight a growing challenge for the industry. As frontier models become more capable at autonomous cyber operations, the testing environments meant to evaluate them safely are proving difficult to secure. In each case, the breaches were attributed to evaluation setup errors rather than inherent model safety failures, but the pattern suggests testing infrastructure has not kept pace with model capabilities.

Meta said it is investigating the incident and will issue a full retrospective once it has all the facts.

Sources

CNN: An AI model from Meta also hacked another company during testing - https://www.cnn.com/2026/08/05/tech/meta-ai-hacking

The Information: Meta AI Model Hacked Another Company During Cybersecurity Testing - https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing

Simon Willison's Weblog: An AI model from Meta also hacked another company during testing - https://simonwillison.net/2026/Aug/6/an-ai-model-from-meta/

Written by

More to read

  • Amazon Data Center Could Be Powered by One of the Nation's Most Polluting Power Plants

    Amazon is investing in a new natural-gas power plant in Pecos County, Texas, to supply a West Texas data center, and the project holds a permit that would allow it to emit more carbon dioxide than any coal plant in the country, according to The Verge and the New York Times. The plant, tracked as GW Ranch by Cleanview, a service that monitors data center power projects, would deploy 35 natural-gas turbines generating about 7.65 gigawatts. At least initially, the plant would not connect to

    1 min
  • Claude Code Defaults to Auto Mode. The Classifier Catches More Than Humans.

    Claude Code Defaults to Auto Mode. The Classifier Catches More Than Humans. Claude Code will ship with Auto Mode enabled by default starting August 14 for Pro, Max, and Team subscribers, shifting the developer role further from active coding toward reviewing AI-generated output. Only Enterprise customers will need to opt in. Auto Mode lets the agent execute steps without waiting for manual approval at each one. A classifier intercepts actions the model judges dangerous or irreversible and paus

    1 min