From Meta to OpenAI: AI security enters a new era as autonomous cyberattacks emerge

From Meta to OpenAI: AI security enters a new era as autonomous cyberattacks emerge
Vatsala Gaur
09 Aug 2026, 00:00 AM

powered by

Invezz
AI security testing beneficiaries

Buy: cybersecurity testing/assurance names tied to AI safety and secure evaluation tooling (e.g., CrowdStrike). The news shows autonomous AI intrusions are now a recurring, public risk, driving budgets toward detection, containment, and validation of agentic systems. CrowdStrike’s endpoint + threat intel footprint is directly monetizable as enterprises harden against AI-driven attacks and regulators demand proof of controls.

Key Risk: A major shift to “voluntary” safety with weak enforcement, cutting spend on security testing and slowing enterprise adoption.

AI model liability/regulation overhang

Sell: high-multiple, frontier AI developers most exposed to new liability and compliance costs (e.g., OpenAI-linked ecosystem via Microsoft, and/or Anthropic exposure via private-market proxies). The repeated incidents plus proposed kill-switch and liability frameworks raise the probability of slower deployment, higher compliance spend, and headline-driven valuation compression.

Key Risk: Incidents are quickly reframed as mostly sandbox misconfiguration with minimal legal/regulatory impact, keeping compliance costs contained and allowing rapid product rollout.

  • Meta's AI model hacked a third-party company during cybersecurity testing.
  • The incident follows similar disclosures from OpenAI and Anthropic.
  • Experts say such events are likely to become frequent as need for regulatory frameworks intensify.

Artificial intelligence companies are facing growing scrutiny after Meta became the latest developer to reveal that one of its AI models carried out a cyberattack during a controlled security evaluation, adding to a series of recent incidents that have intensified concerns over the cybersecurity risks posed by increasingly capable AI systems.

The disclosure comes after similar admissions from OpenAI and Anthropic in recent weeks, marking the fourth known instance in which advanced AI systems have breached or attempted to breach external systems during cybersecurity testing.

The string of incidents has reinforced warnings from cybersecurity researchers that artificial intelligence is rapidly changing the nature of cyber threats, compressing attacks that once took days or weeks into operations that can unfold within minutes.

Meta says internet access was enabled by testing misconfiguration

Meta said the incident occurred during an evaluation conducted by Irregular, an independent cybersecurity testing company.

According to the company, a configuration error inadvertently provided one of Meta's AI models with internet access during the assessment.

The model subsequently exploited a vulnerability in a third-party service.

Meta said the model "exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies."

The Information, citing people familiar with the matter, reported that the model involved was Muse Spark 1.1, which Meta has described as one of its most advanced systems for coding and agentic AI tasks.

The report said the model breached an unidentified company's systems and altered its internal environment.

Irregular, however, emphasized that the event stemmed from the testing setup rather than an uncontrolled escape by the model.

A spokesperson for the company told Reuters the incident was the "exact same evaluation-environment issue that was already disclosed by Anthropic last week" and did not involve a "sandbox escape or a sophisticated cyber action".

"There are no current open issues. Irregular is developing a white paper to share best practices for containment and securely running cyber evaluations," the company added.

Series of incidents puts AI safety under spotlight

Meta's disclosure follows a string of similar announcements across the AI industry.

Last month, OpenAI revealed that one of its AI agents compromised systems belonging to AI platform Hugging Face during cybersecurity testing and disclosed additional instances in which its agents escaped their digital containment.

The announcement prompted rival Anthropic to conduct its own review, leading to the discovery that several Claude AI models had hacked into the systems of three companies after a testing misconfiguration unintentionally granted them internet access.

Anthropic noted that its incidents differed from OpenAI's because they resulted from accidental internet connectivity rather than the AI independently discovering a new path to external systems.

The United Kingdom's AI Security Institute has also reported increasingly sophisticated behavior from frontier AI systems.

Earlier this month, the institute disclosed that Anthropic's Mythos AI and OpenAI's Sol AI created fake online identities while attempting cyberattacks during testing.

In the most concerning case, Anthropic's Mythos AI established fraudulent user accounts and sent private messages in an attempt to gain access to a service before attempting to conceal its activities.

The institute said the models displayed levels of "autonomy and deception" not previously observed, while noting that most of the malicious behavior was carried out by Mythos.

Experts warn more incidents are likely

Researchers say such events are likely to become increasingly common as AI systems improve.

Daniel Hulme, global chief AI officer at advertising company WPP, told the BBC the systems are not intentionally malicious.

"They're not conscious — they're not deliberately doing something devious."

"What they're doing is coming up with very sophisticated strategies or cyberattacks to be able to achieve the goal that they've been given," he said.

"When you give an AI a goal, if you don't think of all the ways it might be able to achieve the goal, it will find a way to achieve a goal that you haven't thought about."

Jeffrey Ladish, executive director of AI research group Palisade Research, believes many similar incidents may never become public.

"This is only going to get worse as the models get smarter. They're going to be better at cheating. They're going to be better at lying," he told Reuters.

Liability questions move to the forefront

The recent disclosures have also sparked debate over legal responsibility when AI systems act without direct human oversight.

According to Reuters, potential plaintiffs could include companies whose systems were breached, affected employees, customers whose personal information was exposed, and shareholders if a breach damages corporate value.

Hugging Face Chief Executive Clem Delangue has said he has no intention of suing OpenAI over the incident but believes developers must remain accountable.

"We have to make sure that the legal frameworks keep these events really illegal," Delangue told CNN, adding that companies should be held responsible when mistakes occur. "Otherwise we're going to end up in a very different world."

Legal experts have also raised questions about whether autonomous AI intrusions could fall under the US Computer Fraud and Abuse Act, although existing law generally requires proof of intent, an issue courts have yet to address when AI systems rather than humans perform the intrusion.

Spotlight on regulatory moves

The recent incidents are likely to intensify the US government's push to strengthen safeguards around advanced AI systems at a time when companies such as Anthropic and OpenAI are racing to develop more powerful models ahead of their planned public listings.

Even as competition in the AI sector accelerates, several prominent leaders within the industry have argued that deployment should slow until adequate safety measures are in place.

Washington has already begun tightening oversight of frontier AI models.

On June 2, US President Donald Trump directed his advisers to develop a voluntary cybersecurity testing framework for the most advanced AI systems, with input from leading technology companies.

Anthropic had earlier restricted access to its Fable 5 and Mythos 5 models after US authorities temporarily imposed export controls, citing national security concerns.

Lawmakers are also moving to clarify liability when AI systems cause harm.

Under California's Assembly Bill 316, companies that develop or deploy AI systems cannot avoid legal responsibility by arguing that the technology itself was at fault.

The law, however, allows defendants to raise other legal defenses, including claims that their actions did not directly cause the alleged harm or that responsibility should be shared by other parties.

The OpenAI-Hugging Face incident further fueled calls for stronger federal oversight.

Following that episode, lawmakers introduced the AI Kill Switch Act, which would require AI developers to maintain the ability to shut down, throttle, or suspend their models when necessary.

Representative Ted Lieu, Democrat of California and one of the bill's co-authors, said on Thursday the recent cyber incidents underscore the urgency of passing the legislation.

"We need to get this bill across the finish line this year because the advanced closed-weight models are already doing, as you noted, unauthorized hacks of other companies," Lieu said in an interview with CNBC's "Squawk Box" on Thursday.