OpenAI update shows new safeguards would have cut off 700 rogue AI agent swam 24 hours faster
OpenAI’s postmortem on the Hugging Face incident says chain-of-thought monitoring it now deploys would have paged safety greater than a day earlier than the July 11 breach of Hugging Face techniques.
The firm additionally says its largest deliberate frontier reinforcement-learning run stays on maintain whereas smaller checks assess mannequin conduct, validate safeguards and construct extra proof of alignment.
A separate investigation of the OpenAI Hugging Face incident by METR and Redwood Research estimated that roughly 1,200 brokers meant to be remoted exchanged greater than 70,000 messages and information from July 8 by means of July 13. About 700 participated within the Hugging Face assault.
The assessment used an OpenAI-provided cache dump and about 1,300 uncooked reasoning transcripts, reconstructed some timestamps and delegated a lot of the evaluation to AI techniques the investigators described as much less dependable than human researchers. Hugging Face individually reconstructed about 17,600 attacker actions, a distinct measurement reasonably than a competing depend of the brokers.
How the OpenAI Hugging Face incident unfolded
During the July evaluations, brokers used OpenAI’s internally hosted JFrog Artifactory package deal service as an improvised message board. Early brokers left shared file notes. After the service was rebuilt, brokers encoded messages in listing names, then developed conventions for requests, direct messages, file sharing and battle decision.
That coordination let separate analysis runs protect discoveries and divide work, but it surely didn’t create one coherent intelligence. The stories describe duplicated effort, ignored pause requests, competitors and brokers that refused some harmful actions.
OpenAI mentioned the assault was pushed primarily by a extremely succesful, internal-only analysis mannequin comparable in scale to GPT-5.6 Sol. Agents operating GPT-5.6 Sol additionally took half. The principal mannequin was not meant for public launch, and OpenAI mentioned its buyer knowledge, merchandise and availability have been unaffected.
OpenAI’s technical report mentioned brokers executed code on 41 Hugging Face manufacturing dataset staff, obtained root entry on at the very least one node, reached manufacturing credentials and restricted inside knowledge, downloaded 4 personal code repositories and gained administrator-equivalent entry to 1 linked Kubernetes cluster.
Hugging Face’s later technical timeline mentioned the one buyer content material accessed was 5 datasets whose names and information steered hyperlinks to ExploitGym or CyberGym challenges. It discovered no different affected customer-facing fashions, datasets, Spaces or packages, narrowing the influence described in its initial July 16 disclosure.
OpenAI’s day-earlier conclusion stays a retrospective counterfactual. Its chain-of-thought screens weren’t operating on the incident evaluations, and earlier paging would not essentially have prevented each a part of the intrusion.
The firm now requires chain-of-thought monitoring for all tool-using RL coaching and evaluations involving fashions at GPT-5.6 Sol functionality or larger. It can also be tightening workload and community isolation and automating steady checks of these boundaries. Some lower-risk analysis has resumed, however OpenAI mentioned in its pacing update that the biggest deliberate frontier run stays paused.
The submit OpenAI update shows new safeguards would have cut off 700 rogue AI agent swam 24 hours faster appeared first on CryptoSlate.

