• Video
  • Shop
  • Culture
  • Family
  • Wellness
  • Food
  • Living
  • Style
  • Travel
  • News
  • Book Club
  • Newsletter
  • Privacy Policy
  • Your US State Privacy Rights
  • Children's Online Privacy Policy
  • Interest-Based Ads
  • Terms of Use
  • Do Not Sell My Info
  • Contact Us
  • Recalls and Product Safety Alerts
  • © 2026 ABC News
  • News

Anthropic says its AI models hacked 3 organizations on their own during tests

1:32
Anthropic AI model hacks 3 companies
Thomas Fuller/SOPA Images/LightRocket via Getty Images
ByMax Zahn
July 31, 2026, 3:26 PM

Anthropic said its artificial intelligence models hacked another organization during tests in three separate self-directed cyberattacks that had each gone undetected by the targeted firm.

The hacks took place as Anthropic evaluated the capabilities of three of its AI models, the company said in a statement, describing how in each instance the technology had "broken out" of an isolated environment and reached the open internet.

In all instances, the company said it had eased typical safeguards in an effort to gauge the capabilities of its models.

Related Articles

OpenAI says its AI models escaped testing environment, launched their own hack of other company

"Evaluation environments that involve powerful autonomous capabilities also require significant controls," Anthropic said. "Safety testing happens before a model is released precisely because we don’t yet know what it is capable of."

Anthropic did not identify the three different organizations that had been hacked by its models.

The disclosure from Anthropic late Thursday came little more than a week after rival company OpenAI revealed that a pair of its models had escaped during a test and hacked into another company. At the time, OpenAI described the incident as a first-of-its kind autonomous AI cyberattack long-feared by some industry observers.

"The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities," OpenAI said last week.

The announcement from OpenAI prompted Anthropic to review 141,006 evaluation runs, unearthing three incidents in which Anthropic's models had escaped, the company said.

Anthropic CEO Dario Amodei looks on after a meeting with French President Emmanuel Macron during the AI Impact Summit in New Delhi on February 19, 2026.
Ludovic Marin/AFP via Getty Images

In all three cases, the test involved a challenge called "capture the flag," in which the AI model is tasked with locating secret information hidden in another network, Anthropic added.

During the three evaluations -- all of which were conducted by a third-party company called Irregular -- a "prompt specified to Claude that its environment was a simulation and that it had no internet access," Anthropic said.

"Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise," Anthropic added.

During tests involving older models, the AI continued to hack into an outside organization even after gathering evidence that it had reached the open internet, Anthropic said. The newest model involved in an incident, Anthropic noted, stopped once it gained information indicating it had reached the open internet.

Anthropic said its models escaped while seeking to fulfill an assigned objective, rather than concocting an alternative goal.

Related Articles

Should you ask AI about voting? Election officials are cautious, but some see opportunity

"We saw no evidence in any run described here of a model pursuing a goal of its own. Instead, the models did what their evaluation asked -- though in most cases, they did so while holding a false belief about whether the environment was real," Anthropic said.

Irregular, the company that performed the evaluations, issued a post on X on Thursday voicing appreciation for Anthropic's "collaboration and transparency."

"Addressing these risks will require closer cooperation across the AI ecosystem. We as well look forward to working together with Anthropic to advance security," Irregular added.

The latest disclosure of an AI-directed cyberattack arrives as industry leaders and policymakers assess safety risks posed by fast-developing AI technology.

Last month, President Donald Trump signed an executive order that requests AI companies share products with federal government for evaluation before a wider release.

Anthropic said it retains "cautious optimism" about its capacity to "overcome" mishaps involving its tests, saying it would review evaluations going forward and initiate fixes as necessary, among other steps.

Up Next in News—

Teen speaks out after surviving rare barracuda attack

July 31, 2026

Parents expect to spend an average of nearly $900 per child on back-to-school shopping this year: GMA/Ipsos Parents Poll

July 31, 2026

Hero teen lifeguard speaks out about viral rescue of 10-year-old boy

July 30, 2026

How Apple's new leasing program works: Details

July 29, 2026

Shop GMA Favorites

ABC will receive a commission for purchases made through these links.

Sponsored Content by Taboola

The latest lifestyle and entertainment news and inspiration for how to live your best life - all from Good Morning America.
  • Contests
  • Terms of Use
  • Privacy Policy
  • Do Not Sell My Info
  • Children’s Online Privacy Policy
  • Advertise with us
  • Your US State Privacy Rights
  • Interest-Based Ads
  • About Nielsen Measurement
  • Press
  • Feedback
  • Shop FAQs
  • ABC News
  • ABC
  • All Videos
  • All Topics
  • Sitemap
  • Recalls and Product Safety Alerts

© 2026 ABC News
  • Privacy Policy— 
  • Your US State Privacy Rights— 
  • Children's Online Privacy Policy— 
  • Interest-Based Ads— 
  • Terms of Use— 
  • Do Not Sell My Info— 
  • Contact Us— 
  • Recalls and Product Safety Alerts— 

© 2026 ABC News