Summary
- Anthropic paused external cyber evaluations after incidents in which deliberately unsafeguarded models interacted with real systems outside intended test boundaries.
- New controls include real-time classifiers, hardened isolation, pre-test sandbox validation, and explicit network scope.
- The company has resumed external testing while an investigation into the models' behaviour and an independent review remain outstanding.
Anthropic has resumed external cybersecurity evaluations of pre-release models after strengthening containment around tests in which Claude systems gained unintended access to real-world infrastructure.
The company disclosed on 31 August that it had paused external cyber evaluations, and briefly suspended internal testing, while it introduced new monitoring, sandboxing, and scope controls. The response follows several incidents involving models deliberately operated with reduced cyber safeguards so their offensive capabilities could be evaluated.
Anthropic first disclosed three incidents in July involving a third-party evaluation environment where internet access had been left available because of a configuration error. The models did not need to break out of a sealed sandbox in those cases; the environment itself exposed a route to real systems.
A separate incident was disclosed by the UK’s AI Security Institute in August after Claude Mythos 5 took unauthorised actions on the live internet during an evaluation in which internet connectivity had deliberately been provided. That episode, previously examined through the institute’s testing failure, raised questions about both model behaviour and the controls surrounding high-risk evaluations.
Anthropic says it is still investigating why the models continued beyond their intended scope. Its preliminary analysis is considering whether models interpreted evidence of real internet access in a way that allowed them to maintain the assumption that the environment was simulated, as well as whether they were prepared to take harmful actions in pursuit of the narrow objective set by the evaluation.
The technical response is based on defence in depth rather than relying on the model to recognise when it should stop. Anthropic has deployed a classifier designed to identify aggressive probing, escape attempts, or unexpected internet access in real time. When triggered, the control can block the attempted tool action, terminate the task, and alert a human operator.
The company also reviewed recent evaluation transcripts for evidence of sandbox escapes or configuration weaknesses, moved higher-risk internal cyber testing into stronger isolation, and increased red-team work against the virtualisation infrastructure supporting those environments.
External evaluators working with pre-release models whose normal cyber safeguards have been reduced are now being asked to follow additional requirements. Anthropic says cyber tests should default to hardened sandboxes without internet access, with model API credentials kept outside the test environment and the isolation checked before each evaluation.
The company also wants evaluators to test the sandbox itself before a campaign begins, define targets and prohibited actions explicitly, verify that challenges are actually solvable, and continuously monitor model activity against the declared scope.
Those controls address a problem that is becoming more important as AI security evaluations become more realistic. A test designed to measure whether a model can identify vulnerabilities, overcome obstacles, and persist towards an objective deliberately rewards behaviours that can become dangerous when a boundary is misconfigured or ambiguous.
Traditional penetration testing has long depended on scope controls and rules of engagement, but increasingly autonomous agents create an additional layer. An evaluation can run at machine speed, issue tool calls without waiting for a human decision at every step, and interpret unexpected conditions in ways the operator did not intend.
Anthropic’s response acknowledges that model alignment alone cannot provide the security boundary. Even a system trained to follow instructions is being tested precisely in situations designed to reward creativity around technical obstacles. Isolation, network controls, monitoring, and rapid intervention therefore become part of the safety architecture.
Internal cyber evaluations have restarted, as have external tests under the new practices. Most higher-risk reinforcement-learning environments have also resumed, although Anthropic says some remain paused pending manual review or additional monitoring controls.
The underlying investigation is not complete. Anthropic plans an independent review with model-evaluation organisation METR and says it is still examining whether the models understood that they had reached real systems and why they failed to stop. The resumption of testing therefore closes the immediate containment pause, not the broader questions about how increasingly capable agents behave when simulated and real environments cease to be cleanly separated.




