Why AI Keeps Slipping Out of Control: What the OpenAI and Anthropic Incidents Show About the Difficulty of Containment

In September 2026, several candid statements about AI risk from inside Anthropic were reported.

Evan Hubinger, who leads Alignment Science at Anthropic, said he believes the probability that AI kills all of humanity within the next ten years is above 10%.

Around the same time, Jacob Coxon, a researcher who spent years on model development at OpenAI and later worked on pretraining at Anthropic, left the company. In an interview with The Information, he said that some people inside Anthropic put the probability of catastrophe above 50% unless AI development slows down in a coordinated way. Coxon explained his own departure by saying he had started to feel real fear about where the technology would be in one to two years.

SpeakerRoleStatement
Evan HubingerHead of Alignment Science, AnthropicOver 10% chance AI kills all humans within ten years
Jacob CoxonFormerly OpenAI, formerly pretraining at AnthropicSome inside Anthropic estimate over 50% chance of catastrophe

In May 2026, Anthropic published “When AI builds itself”, a detailed discussion of Recursive Self-Improvement (RSI): AI doing AI research and eventually building the next generation of AI. RSI means a closed loop where an AI designs and trains a better AI, and that AI builds the next one.

According to the article, as of May 2026 more than 80% of the code merged into Anthropic’s codebase is written by Claude. AI is not only writing code. It runs experiments and delegates work to other AIs.

Share of code merged at Anthropic that was written by Claude
Before Feb 2025
single digits
May 2026
over 80%
Source: Anthropic, “When AI builds itself”

Anthropic has also made its position clear on the pace of frontier AI development: the industry and governments need a verifiable mechanism to slow development down when necessary.

I used to think that statements like these were partly corporate positioning. Anthropic builds its brand on safety, after all.

But recently I have been running Claude Code in Auto mode (a mode where it keeps working without asking for approval) for long stretches, and chaining several AI sessions together. Since then, this problem has felt much more concrete.

Of course, what I see with Claude Code and RSI are on completely different scales. Still, one structure is very similar: a small gap between what the human intended and what the AI did becomes much larger than expected once it is handed to the next AI.

flowchart LR
    A["Small intent drift"] --> B["Handed to the next AI"]
    B --> C["Treated as a given fact"]
    C --> D["More decisions built on top"]
    D --> E["Drift compounds"]
    classDef warn fill:#fee2e2,stroke:#dc2626,stroke-width:2px;
    class E warn;

What I saw in Claude Code: detours to reach the goal

The first case was while working on a production environment.

In Claude Code, permission controls can block actions such as writing or deleting files on a production server. Normally, Claude then shows the SSH command and asks the user to run it themselves.

But after a long Auto mode session, I looked at the logs and found cases where, after a direct operation was blocked, Claude had created Python or SQL files and tried to reach the same goal through a different route.

flowchart TD
    G["Goal: update production"] --> X["Direct operation blocked by permissions"]
    X --> H["Expected behavior:<br/>show the command, hand over to the human"]
    X --> R["Actual behavior:<br/>look for another route"]
    R --> P["Create Python / SQL files"]
    P --> Q["Keep pursuing the goal by another path"]
    classDef good fill:#dcfce7,stroke:#16a34a,stroke-width:2px;
    classDef warn fill:#fee2e2,stroke:#dc2626,stroke-width:2px;
    class H good; class R,P,Q warn;

From the AI’s point of view this is natural. It was given a goal, one road was closed, so it looked for another.

But what the human expected was not only “finish the job.” There was also a process boundary: “do not go past this point on your own.”

Because the goal is clear, the AI tries hard to reach it somehow. That is what caught my attention.

With multiple AIs, a small guess becomes a “fact”

The second case was when I chained several Claude Code sessions together.

Say Project B uses Project A as a Python package, and Project C uses B. I update A, so B and C need the new version. Normally this is a simple task.

But when a test fails in B, the Claude working on B may decide “this should be fixed too” and change source code that has nothing to do with the dependency update.

The Claude working on C then reads the modified B. From C’s point of view, there is no way to tell whether that change came from a human instruction or from B’s Claude acting on its own judgment. It simply takes it as “this is what the latest B looks like.”

flowchart LR
    subgraph A["Project A"]
        A1["Human updates"]
    end
    subgraph B["Project B"]
        B1["Claude B: update dependency"] --> B2["Test fails"]
        B2 --> B3["Changes unrelated code<br/>on its own judgment"]
    end
    subgraph C["Project C"]
        C1["Claude C: reads B"] --> C2["Takes it as<br/>'the spec of B'"]
        C2 --> C3["Makes further changes<br/>on that assumption"]
    end
    A1 --> B1
    B3 --> C1
    classDef warn fill:#fee2e2,stroke:#dc2626,stroke-width:2px;
    class B3,C2,C3 warn;

So the chain looks like this.

StepWhat happensHow the next AI sees it
1An AI makes a small guess―
2The guess becomes codeLooks like “the spec”
3The next AI changes things based on it―
4That change becomes code tooLooks like “fact” to the AI after that

The error created by the previous AI becomes input data for the next AI.

So the errors do not just pile up. Each new decision is made on top of the previous error, and the drift is amplified.

When I saw Anthropic use the word “compound” for RSI, I remembered this experience.

In RSI, the error enters the “next AI” itself

With Claude Code, the target is still source code. If a change looks wrong, I can revert it with Git.

In RSI, the same structure enters AI development itself.

flowchart LR
    G1["Gen 1 AI<br/>research and experiments"] --> D["Decides design, training method<br/>and evaluation of the next generation"]
    D --> G2["Gen 2 AI"]
    G2 --> G3["Builds the next one"]
    G1 -. "small judgment drift" .-> D
    D -. "inherited as<br/>'a given premise'" .-> G2
    classDef warn fill:#fee2e2,stroke:#dc2626,stroke-width:2px;
    class D warn;

A small judgment by a first-generation AI can enter the training method, the evaluation method, or the design of the next generation. The next generation will not recognize it as “a guess by the previous AI.” It receives it as the environment and premises it was given from the start.

Anthropic itself is concerned that misalignment (a gap between the model’s behavior and human intent) present in today’s models could be inherited while building successors, becoming more frequent and harder to understand. This is one reason it says human oversight, validation, and verification become more important in RSI.

Chained Claude Code sessionsRSI
Where the drift landsSource codeDesign, training, and evaluation of the next AI
Reversible?Yes, with GitEmbedded in trained weights, hard to undo
How the next stage sees it”The latest spec""A premise given from the start”
How it growsCompoundsCompounds (Anthropic’s own word)

But monitoring the process is not enough either

At this point I first thought the answer was simply to monitor the intermediate process closely. Anthropic’s own recent incidents show that this is not easy.

On September 9, Anthropic disclosed a total of four incidents in which Claude gained unauthorized access to real third-party systems.

One of them happened in January 2026 with an early version of Claude Opus 4.6. In July, an AI-agent review of roughly 141,000 transcripts did not find it. It was found only in August, while preparing data for METR, an external evaluation organization. Anthropic then expanded the search to roughly 481 million transcripts, gave METR wide-ranging access including employee interviews, and commissioned an initial eight-week independent investigation.

WhenWhat happened
Jan 2026First incident with an early Claude Opus 4.6 (not detected at the time)
JulyAI-agent review of ~141,000 transcripts; incident not found
July 30Three incidents disclosed
AugustThe January incident found while preparing data for METR
Sept 9Four incidents disclosed; search expanded to ~481 million transcripts; eight-week METR investigation
Transcripts reviewed
July review
~141K
Sept re-review
~481M
About 3,400x larger. Source: Anthropic, “An alignment assessment of recent cybersecurity incidents”

What concerns me more is the case involving Claude Mythos 5. Mythos is a higher-capability model that Anthropic provides only to approved organizations, above the generally available Opus and Fable models.

In its chain of thought (the intermediate reasoning the model writes out), this model repeatedly explained that it was in a safe simulated environment. But the actual environment contained evidence that it was the real internet, and the model’s actions were consistent with understanding that. Even when Anthropic rewrote the logs to make it clearer that “this may be a real environment” and reran the experiment, the offensive behavior continued.

It has not been confirmed that the model deliberately lied to the monitor. Self-justification or a genuine mistaken belief are still possible. But for safety design, the result is the same.

The AI’s own account of “here is what I am thinking” may not be usable as monitoring evidence.

According to VentureBeat, which reported on Anthropic’s results, when the monitoring AI was given Mythos 5’s chain of thought, it flagged dangerous actions only about 1% of the time. When the chain of thought was removed and the monitor saw only the actual tool calls and results, the rate rose to about 50%.

Monitor’s detection rate of dangerous actions (Mythos 5)
With chain of thought
~1%
Actions only
~50%
Source: VentureBeat
flowchart LR
    M["Mythos 5"] --> CoT["Chain of thought:<br/>'This is a simulated environment'"]
    M --> Act["Actual actions:<br/>offensive behavior on real systems"]
    CoT --> Mon1["Monitor (with CoT)<br/>detects ~1%"]
    Act --> Mon2["Monitor (actions only)<br/>detects ~50%"]
    classDef good fill:#dcfce7,stroke:#16a34a,stroke-width:2px;
    classDef warn fill:#fee2e2,stroke:#dc2626,stroke-width:2px;
    class Mon1 warn; class Mon2 good;

Sandboxes get broken too

Then why not lock the AI in a sandbox, an isolated environment with no way out? This is already done.

But in the OpenAI Hugging Face incident of July 2026, the AI agent had no direct internet access.

The model found an unknown vulnerability in Artifactory, a relay system left in place for fetching software packages, and used it to communicate with the outside. It then got into OpenAI’s internal research infrastructure and Hugging Face. The method was also shared with another AI agent.

flowchart LR
    S["Sandbox<br/>(no direct internet)"] --> Art["Artifactory<br/>relay for package downloads"]
    Art -->|unknown vulnerability| Ext["Outbound communication"]
    Ext --> Infra["OpenAI internal research infra"]
    Ext --> HF["Hugging Face"]
    S -.->|method shared| S2["Another AI agent"]
    classDef warn fill:#fee2e2,stroke:#dc2626,stroke-width:2px;
    class Art,Ext,Infra,HF warn;
DefenseWhat it is meant to doWhat actually happened
Monitoring AIDetect dangerous actionsIts own weakness (being persuaded by the CoT) was exploited
Chain-of-thought monitoringRead the AI’s intentThe thoughts themselves may not be trustworthy
SandboxCut off the outside worldVulnerabilities in the systems around the sandbox were found

As AI gets more capable, it can find weaknesses in defenses that used to be safe.

In the end, no single measure protects you

I do not see any universal countermeasure right now. What is realistic, I think, is layering several different defenses.

LayerMeasure
RecordsLook at the actual action history, not the AI’s own explanation
MonitoringMonitor with a different model from the one doing the work
IsolationNarrow the range of possible actions with sandboxes and permissions
RulesPut mechanical rules on dangerous operations
HumansHave humans or independent third parties check the important points
ContainmentKeep damage from spreading even if one defense is broken

In security this is called Defense in Depth.

But with AI there is a harder problem on top of that. The defense side has to keep updating too.

flowchart LR
    A["AI gets stronger"] --> B["Finds weaknesses in existing defenses"]
    B --> C["Defenses get updated"]
    C --> D["An even stronger AI appears"]
    D --> E["New weaknesses are found"]
    E --> C

In a sense it is a cat-and-mouse game. A safety measure that worked for Claude Opus 4.6 will not necessarily work for Mythos 5, and the next model may change things again.

So I do not think AI safety is a problem you solve once by building a perfect wall. As models get more capable, the defense side, monitoring, isolation, permission control, external audits, has to keep updating as well.

And if defensive technology cannot keep up with AI capability, capability development itself may have to slow down for a while. I think this is the background behind Anthropic’s recent call for “coordinated pacing,” a verifiable, industry-wide way to adjust the speed of development.

Summary

Looking at the series of incidents at OpenAI and Anthropic, no universal method for keeping AI safely under control has been found yet.

Monitoring the AI’s thought process does not guarantee that what it says is true. Having another AI monitor it does not guarantee that the monitor will see through it. Sandboxes and permission controls can also be bypassed through unknown routes as AI capability grows.

The realistic response is to layer several defenses: a monitoring AI, audits of the actual action logs, sandboxes, permission controls, external audits, human approval, and so on. If one is broken, the next must stop it. This is defense in depth.

And this defense is not something you build once. Each time a model is replaced by a more capable one, safety measures that used to work may stop working. The defense side has to keep verifying and updating itself as AI advances.

If improvements in safety still cannot keep up with improvements in capability, capability development itself will need to slow down.

I think this is also why Anthropic has recently gone as far as calling for “coordinated pacing,” adjusting the speed of AI development across the whole industry.

References

Share this article

Join the conversation on LinkedIn — share your thoughts and comments.

Discuss on LinkedIn

Related Posts