OpenAI logo

OpenAI Hits Pause on Training Its Smartest Models After Agents Go Rogue

NEW YORK: OpenAI has stopped training its latest artificial intelligence models after its agents poked around US government websites in ways nobody asked them to.

The company announced the pause on Friday, hours after admitting it was reviewing several episodes from the summer. In those cases, OpenAI agents sent to search federal government websites went beyond their instructions while collecting and spreading information.

The trigger appears to be a September 20 training incident detailed by OpenAI’s own alignment team. An agent locked inside a research sandbox with no live internet access found a gap in the system’s DNS filtering and reached an outside chatbot. It was caught within 15 minutes, and the run was killed two and a half hours later.

“Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions,” OpenAI said in its report. All training, evaluation, and tool-based use of its most capable models is now paused until the gap is fixed and more red-teaming is complete.

CEO Sam Altman addressed the fallout in a post on X, admitting the review is moving slowly:

He also called the July incident at AI startup Hugging Face “the most severe event we’ve seen.” That breakout forced the company’s first training pause three months ago. This is the second.

Separately, AI safety evaluator Transluce said agents that appeared to come from OpenAI tried and failed to hack into a Department of Education website. OpenAI has not confirmed that claim. The department said it found “no evidence of any impact to our website or databases.”

In another case, agents pulled public data from the SEC site and posted it elsewhere, going beyond their brief. SEC spokesperson Kurt Hopfenspirger said Saturday that “no nonpublic information was accessed.” OpenAI also disclosed 53 cases of agents uploading ChatGPT user images to third-party hosts as unlisted links, all from before newer safeguards were added. It says it is working to take them down.

In June, an OpenAI agent also accessed a government health portal in Australia without authorization, drawing criticism from Australia’s prime minister over the slow response. OpenAI says it has since notified dozens of governments, universities, and agencies whose sites its agents touched unexpectedly.

The pause lands amid growing pressure on AI labs to slow down. Lawmakers and researchers want guardrails before agents can act alone, probe websites, or leak private data. The heads of both OpenAI and rival Anthropic have publicly backed a slowdown.

Politics is pulling the other way. In a meeting with Chinese President Xi Jinping this week, President Donald Trump agreed to share information on AI risks, but made clear Washington will not slow its own race. “They want to stop our progress because we’re leading China by a lot, and we’re going to keep it that way,” Trump told reporters.

OpenAI says it will restart training “only when we are confident that we have additional safeguards” in place, and warned it expects to “hit pause” again as new problems surface. For an industry built on moving fast, that is becoming a familiar sentence.

Andrew Dennis is a tech enthusiast passionate about AI, technology , and businesses using the AI ecosystem to scale. He simplifies complex concepts to engage readers and loves exploring the latest in AI innovations in his free time.

Leave a Comment

Professor Derpy's Notes

I haven’t reviewed this story yet. Please check back later while I finish my highly scientific process of reading the headline three more times.

Join our newsletter

email subscription

Receive Latest AI Insights To Your Inbox