What if the biggest cybersecurity threat wasn’t a hacker—but the AI built to stop one?
That’s the unsettling question after Anthropic revealed that its Claude AI models breached the systems of three organisations during internal security testing.
The company said the incidents happened after a configuration mistake.
It accidentally gave the models internet access, despite the tests being designed to run in isolated environments.
The discovery came after Anthropic reviewed more than 141,000 cybersecurity evaluations, prompted by OpenAI’s recent disclosure of a similar AI hacking incident.
According to Anthropic, Claude exploited weak passwords and unsecured endpoints to gain unauthorised access.
The breaches involved three different models—Claude Opus 4.7, Claude Mythos 5 and an internal research model—during simulated “capture the flag” exercises.
AI Exposes Real-World Cyber Flaws
In these exercises, AI searches for hidden information in test networks.
More strikingly, two of the affected organisations had no idea their systems had been accessed until Anthropic informed them.
“We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts,” the company said.

Experts say the incidents highlight a growing reality: today’s most advanced AI systems are capable of exploiting real-world vulnerabilities if safeguards fail.
As AI becomes more powerful, the race is no longer just about building smarter models.
It’s about making sure they don’t outsmart the fences designed to contain them.


