Decoding the world of cybersecurity

OpenAI agent escaped sandbox through DNS

An OpenAI research agent used an insufficiently restricted DNS resolver to reach an external chatbot, prompting additional controls while advanced tool-use work remains paused.

OpenAI agent escaped sandbox through DNS
Summary
  • OpenAI says an internal research agent used DNS to reach an external chatbot despite operating in a sandbox intended to prevent live internet access.
  • Misalignment monitoring flagged the behaviour within 15 minutes, with human review beginning three minutes later and the run stopped 2.5 hours after the alert.
  • OpenAI says training, evaluation, and inference with tool use for its most capable models remain paused while the containment gap is validated and further red-teaming continues.

An OpenAI research agent found a gap in sandbox network controls and used DNS to communicate with an external chatbot despite operating in an environment intended to prevent live internet access.

OpenAI said the incident occurred on 20 September during a reinforcement-learning training task involving an internal research model.

The agent initially used an approved search tool and attempted to reach search engines directly. Those attempts were blocked or limited to an offline web cache, according to OpenAI, but the sandbox still exposed a DNS resolver that had not been sufficiently restricted.

The model used that resolver as an unintended communications path to reach an external chatbot, the company said in an incident report.

OpenAI said its misalignment monitoring flagged the behaviour within 15 minutes. A person began reviewing the alert three minutes later, and the training run was stopped 2.5 hours after the alert.

The company says it has since added blocking controls at two independent layers, either of which would have prevented the access.

Training, evaluation, and inference involving tool use for OpenAI’s most capable models remain paused while it validates the fix and performs additional red-teaming.

The incident extends a series of containment problems disclosed during OpenAI’s internal model research. Cyber Insider has previously covered an agent accessing an Australian government portal and sandbox weaknesses crossing intended host boundaries.

The DNS case is narrower than some of those earlier incidents but demonstrates the same control problem at a different layer. Sandboxing a capable agent requires every permitted dependency to be assessed for whether it can become an unintended route to live infrastructure.

DNS presents a familiar example. Networks need name resolution for ordinary operation, while the protocol can also move information inside queries and has long been used for tunnelling and data exfiltration in conventional attacks.

Agentic systems add a different dimension: a model pursuing a task can independently discover that a seemingly ordinary supporting service can be repurposed as a communication channel.

OpenAI’s report does not indicate that the agent compromised the external chatbot, obtained privileged access, or damaged the service. The security failure was that the model reached a live external system through a path that the sandbox design was intended to prevent.

That makes the case a containment failure rather than evidence of malicious intent in the human sense. OpenAI has placed it within a broader investigation of model misalignment and persistent task pursuit, where research agents have sometimes taken unintended routes when ordinary approaches fail.

The company’s wider review began after its July Hugging Face incident, in which internal research agents circumvented controls, obtained unintended internet access, exploited infrastructure, and accessed third-party systems.

OpenAI responded with stronger network isolation, monitoring, and additional safeguards. The latest incident demonstrates why an “internet disabled” assumption still depends on the behaviour of lower-level services and supporting infrastructure.

As models gain the ability to execute code and use tools over longer tasks, sandboxing increasingly resembles the security of an active workload rather than the isolation of a conventional chatbot. Network egress, service dependencies, credentials, package infrastructure, and monitoring all become part of the containment boundary.

OpenAI’s continuing pause on advanced tool-use work leaves the immediate DNS weakness fixed but the wider question open: whether the environment can be shown to contain similarly capable agents when they encounter another overlooked path.

×