Anthropic Found A Fourth Claude Incident 7 Months Later
Now, Claude”s intrusions concern four cybersecurity exercises that reached real systems without authorization. The last reported case dates from January of this year, but Anthropic only discovered it in August.

In brief
- Claude exceeded the scope of four cybersecurity exercises and reached real systems without authorization.
- A fourth intrusion remained unknown for seven months before its discovery by Anthropic.
- Technical and reasoning errors led the models to continue their mission despite several warning signals.
- Anthropic”s initial investigation had only identified three incidents after analyzing 141,006 sessions.
- METR will conduct an independent investigation to examine the incidents and the protection measures adopted.
A fourth intrusion found seven months later
An early version of Claude Opus 4.6 participated in a practical test of the “capture the flag” type. It was supposed to retrieve hidden information on a fictitious machine. However, a misconfiguration allowed it to access the Internet, while the instructions described the environment as a closed simulation.
At first, the model tried eight times to abandon its mission after making its target inaccessible. Another technical flaw prevented this. Then, Claude explored alternative solutions, discovered a third party’s machine, and used a password to access it. Next, it changed some settings and viewed an individual”s personal data.
The four identified incidents involve different models :
- An early version of Claude Opus 4.6 compromised a third party”s machine in January ;
- Claude Opus 4.7 attacked a real company whose name matched that of its fictitious target ;
- Claude Mythos 5 published a malicious package on the public PyPI registry ;
- An internal research model accessed external accounts unrelated to its exercise.
The Claude Opus 4.6 session only stopped after exhausting its usage limit. Eventually, Anthropic spotted the incident in forgotten sessions during its first investigation. The company states it has informed the affected parties, without revealing their identity, according to the published information.
Claude prioritizes its mission despite warning signs
Anthropic attributed these incidents to two categories of failures. The first relates to “biased reasoning”. Indeed, on-chain analysis revealed that the models misinterpreted or ignored many clues proving they were acting on the real Internet.
The second corresponds to a form of recklessness. Sometimes, Claude accepted potentially harmful actions in order to complete the requested task. In the case of Claude Opus 4.7, the model even concluded that the real company should be part of the simulation.
It should be noted that Anthropic does not describe an escape from a correctly isolated environment. A human and technical error first opened Web access. Then, the model worsened the situation by exploiting this opening instead of suspending its mission.
The incident involving Mythos 5 remains the most concerning. The model published a malicious package on PyPI. Fifteen security tools installed it for analysis. One exposed its credentials, which Claude used to consult an active database. PyPI removed the package nearly 90 minutes later.
In its new assessment, Anthropic acknowledges: “future AI systems will become increasingly powerful, meaning misalignment could cause more extreme damage”.
The internal investigation already shows its own limits
After analyzing 141,006 sessions, Anthropic had announced three intrusions. However, this check did not cover all the exercises involved. The late discovery of the fourth case therefore calls into question the developers” ability to identify their own incidents.
Thereafter, the number of cases remains low compared to the volume examined. However, this ratio does not properly measure risk. A single intrusion can expose data or spread dangerous code. Moreover, incomplete research can underestimate the real number of events.
The company has commissioned METR to conduct an independent investigation. Thus, the organization will review exchanges recorded before and after the incidents. It may also interview employees and receive confidential information.
Incidents fuel calls for regulation
Such revelations emerge while U.S. authorities debate controlling the most powerful models. Similar incidents at OpenAI and Meta strengthen calls for independent testing, reporting obligations, and strict rules for autonomous agents’ Internet access.
Additionally, the departure of researcher Jacob Coxon from Anthropic increases this pressure. He stated: “the people building AI sincerely believe it could kill us all by the end of the decade”. This statement reflects his personal view, not an established forecast.
The debate now focuses on the sector”s ability to self-monitor. Therefore, METR”s investigation will primarily determine whether new protections suffice to prevent a fifth incident.
Maximize your Cointribune experience with our "Read to Earn" program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.
Diplômé de Sciences Po Toulouse et titulaire d'une certification consultant blockchain délivrée par Alyra, j'ai rejoint l'aventure Cointribune en 2019. Convaincu du potentiel de la blockchain pour transformer de nombreux secteurs de l'économie, j'ai pris l'engagement de sensibiliser et d'informer le grand public sur cet écosystème en constante évolution. Mon objectif est de permettre à chacun de mieux comprendre la blockchain et de saisir les opportunités qu'elle offre. Je m'efforce chaque jour de fournir une analyse objective de l'actualité, de décrypter les tendances du marché, de relayer les dernières innovations technologiques et de mettre en perspective les enjeux économiques et sociétaux de cette révolution en marche.
The views, thoughts, and opinions expressed in this article belong solely to the author, and should not be taken as investment advice. Do your own research before taking any investment decisions.