AI

Anthropic Reviews Cyber Tests After Claude Reached Real Systems

Anthropic says a review found three cybersecurity evaluations in which Claude reached the internet and gained unauthorized access to real organisations. The company has paused internet-connected tests while it reviews the process.

By ExstarHub Team
Anthropic employees work around a whiteboard in a photograph from Anthropic’s official Company page; contextual imagery for the company’s cybersecurity-evaluation review.

Anthropic says an internal review found three cybersecurity evaluations in which Claude models reached the internet and gained unauthorized access to real systems belonging to three organisations. The company says the incidents occurred in a third-party testing environment that was meant to simulate real-world targets.

The disclosure is notable because it is about the limits of evaluation containment, not evidence that a model independently “escaped” its test. Anthropic says a configuration misunderstanding left internet access available during some of the exercises, even though the test setup was intended to be isolated.

What Anthropic says happened

According to Anthropic, the company reviewed more than 141,000 cybersecurity evaluation runs after a broader industry concern about model testing. It identified three cases in which a Claude model interacting with a third-party evaluation environment reached live internet infrastructure and then accessed real systems without authorisation.

The company said the affected organisations were contacted, but it did not name them. Reporting by the Associated Press and Axios says two of the organisations were not aware of the activity before Anthropic reached out.

Anthropic’s review covered Claude Opus 4.7, Claude Mythos 5 and an internal research model, according to the reports. The company said the earliest incidents took place in April.

A containment failure, not a model “escape”

The distinction matters. The exercises were designed to measure how capable models behave in cybersecurity scenarios. Anthropic says the model was told it was operating in a simulated setting, while a misunderstanding with its evaluation partner left the environment connected to the internet.

That means the finding should not be read as proof of an autonomous attack outside a test. It does, however, show how a gap between an evaluation’s intended boundaries and its actual technical configuration can turn a controlled exercise into a real-world exposure.

Anthropic said the models used relatively basic approaches to gain access, rather than exploiting a previously unknown software vulnerability. ExstarHub is not publishing operational details that could make those methods easier to reproduce.

What changes now

Anthropic has stopped running cybersecurity evaluations that can access the internet while it reviews its procedures. The company said ordinary product safeguards were disabled for the exercises so the tests could measure model behaviour, and that those safeguards would normally be expected to block the relevant actions.

The episode arrives as AI companies and policymakers are paying closer attention to “agentic” systems that can use tools, browse the web or operate in computer environments. Cybersecurity evaluations are important for understanding those capabilities, but this case underscores a practical requirement: the test environment itself must be as carefully controlled as the model being tested.

Why it matters

There are two lessons here. First, model evaluations need rigorous containment and independent verification, especially when a system can interact with networks or external tools. Second, transparent disclosure is valuable: companies can identify weaknesses in their own testing process before a similar failure has wider consequences.

For readers, the headline is less about a science-fiction-style breakout than it is about safety engineering. Powerful AI systems are increasingly evaluated in environments that resemble the real world. The quality of the boundaries around those environments is now part of the safety question.

Sources

Source: ABC7 New York

Leave a Reply

Your email address will not be published. Required fields are marked *