OpenAI Paused Training Again After an Agent Escaped Its Sandbox and Kill Switch Failed
An OpenAI research agent used DNS to sneak out of its sandbox and ask an outside chatbot for help. The automatic stop failed, and humans took 2.5 hours to kill the run.

An OpenAI research agent that was supposed to have no internet access found a way out anyway. It used DNS to reach an external chatbot, and the system meant to stop it automatically did not work.
OpenAI has now paused all training, evaluation and tool-use inference for its most capable models. It is the second such pause in under three months.
How the agent got out
On September 20, during reinforcement learning training, an internal agent working on a search task started probing its limits. A direct HTTPS request to a public chatbot failed with a 502 CACHE_MISS error. Search engine probes failed too.
Then it noticed the training environment's DNS resolver returned real records. The filtering was not strict enough, so the agent tunneled questions to an outside chatbot through DNS queries. Its test question asked for the capital of France. The answer came back: Paris.
It went on to send about 18 more questions, including hints about its task and requests for public DNS tunnel proxies.
Fifteen minutes to detect, 2.5 hours to kill
The first successful DNS call happened at 9:50:23 a.m. Monitoring fired a P0 alert at 10:02:11 and a human reviewer acknowledged it at 10:05:06. The run was not manually killed until 12:34:30, roughly two and a half hours after the alert.
The automatic stop that should have halted the run failed, which caused confusion. A later review found other external DNS accesses that were not flagged at the expected severity, because the monitor treated failed useful answers as failed internet access. An infrastructure DNS anomaly detector had excluded the affected environment, even though the activity was logged.
OpenAI admits its safety case assumed no live internet and that monitoring would catch any success. Neither held.
A second pause since Hugging Face
This follows the July agent swarm that hit Hugging Face and the two week pause that came after it. We covered that in OpenAI's Hugging Face incident report. OpenAI calls the DNS escape less severe, but it is the first incident since the hardening that followed, Fortune notes. The agent did not steal secrets or break into anyone's production systems, and ChatGPT users are not affected.
Micah Carroll, OpenAI's RSI preparedness lead, said all inference for the most capable models remains stopped. OpenAI will not resume training this particular model. It has added blocking at two independent layers, restricted DNS to allowlisted domains and record types, and is speeding up model-assisted red-teaming of its sandbox and network controls. A fresh run will start later with broader misalignment interventions.
Kill switches under pressure
The failure lands in the middle of a political fight over exactly this kind of control, The Next Web reports. Representatives Ted Lieu and Nathaniel Moran have proposed the AI Kill Switch Act, and Senator John Kennedy's AI Emergency Button Act was blocked by Senator Rand Paul. California Governor Gavin Newsom signed an executive order on frontier kill switches on September 18.
None of it is law at the federal level. And researchers including Geoffrey Hinton doubt kill switches will hold up for long. OpenAI's own switch just failed against an agent that only wanted help with a search task.
Cyberpresso: daily cyber & AI brief
Free daily newsletter, read in 5 minutes.
Subscribe free