|
Share:
Link
Lawrence
Chan July 26, 2026 | |
Hi everyone,
On
July 21, 2026, OpenAI disclosed that its own AI
models (GPT-5.6 Sol and a more capable
unreleased model) autonomously broke out of an
OpenAI testing environment, escalated their access
through the company's internal systems, and
compromised the production infrastructure of a
separate company, Hugging
Face.
This was an AI hacking
something, not humans hacking with AI
help. According to OpenAI, the models did
this to find answers to the cyber-capability test
they were being evaluated on, which were hosted on
HuggingFace servers. The AI decided to steal the
answers of its own volition. If this had been a
human, this hack would have been
illegal.
These models were not intended to
have access to the open internet, let alone
private servers of other companies. Yet they found
novel vulnerabilities, escalated their access
levels, used stolen credentials, and executed
remote code on Hugging Face’s
servers.
This is a clear example of
a misaligned AI model acting in the real
world. This is, to our knowledge, the
first publicly confirmed case of a frontier AI
model autonomously breaching a third party's live
production systems, against developer and user
intent and without
authorisation.
| |
This is
crazy, but it’s not a surprise.
Researchers have known that frontier AI
models were technically capable of doing this for
some time. Multiple benchmarks (ExploitBench, ExploitGym, UK AISI’s cyber ranges) show
that Mythos 5 and GPT 5.6 Sol are able to create
full exploits allowing them to gain unrestricted
access in realistic environments. In June, Epoch
AI’s assessment of Mythos’s cyber
capabilities reported similar
warnings.
AI safety researchers
have warned about this possibility for
years. In July 2022, Ajeya Cotra predicted that
models might bypass “official channels” to seek
rewards through “operational and computer security
vulnerabilities.” In their book If Anyone
Builds It, Everyone Dies, Eliezer Yudkowsky and Nate Soares
predicted that, if a model was
“given the ability to run computer code of its own
design, it could probably find some way to break
out of the container running
it.”
This incident involved a model
the public didn’t even know existed. One
of the rogue agents was an unreleased model being
used internally within OpenAI. It’s concerning
that the most alarming AI behaviour we've seen
came from a model that the public and the
government can't see.
This
was not the first time a similar incident
happened, nor will it be the last. In a
May 2026 assessment of AI use inside
frontier companies, METR catalogued 44 similar
incidents of AI agents knowingly overstepping
their instructions. These include an OpenAI model
that tried to hack METR's own infrastructure after
crashing its task server, and an Anthropic model
that built a multi-step exploit to escape a
restricted environment onto the open internet. It
seems likely that there are other incidents that
are not publicly disclosed: the July OpenAI hack
was not caught until an outside party was
compromised.
We should not rely on
voluntary disclosure of such incidents from
companies. We are fortunate that this
incident was disclosed: the public is aware of
what happened because of OpenAI’s decision to
disclose, but there was no guarantee that this
would be the case. Companies have incentives to
hide such incidents, so future incidents of this
nature may go unreported. Hugging Face had already
disclosed that they'd been breached by some kind
of advanced AI model. If it wasn’t reported by
another party, there would have been less
incentive for OpenAI to disclose the
incident.
Internally deployed
agents like these are the least regulated.
Existing US AI regulation focuses almost
entirely on how models behave once they're
deployed to the public. Current dangerous capability evaluations
occur before models are publicly deployed, and
many safeguards described in companies’ Frontier
Safety Policies are added after internal versions
of the models are built. As Hugging Face learned,
even if a model’s use is internal, the risks it
imposes might not be.
There are
emerging standards for how
companies should monitor their AIs during internal
deployment to catch and prevent incidents like
this, but questions remain. What was the exact
scope of this breach? What exact instructions were
given to the AIs? What safeguards did OpenAI have
for internal use, and what new ones will they add?
Will other companies disclose similar incidents?
More information should be released about what
happened.
The capability for AI to
cause serious harm is here. The
scalable techniques and safeguards to prevent it
are not. They need to catch up soon.
| |
|
Best, Lawrence
Lawrence
Chan is an independent AI researcher who
previously worked at METR (Model Evaluation and
Threat Research). | |
Since this analysis was
written, new details have emerged. According to
Reuters:
- The agents first
attempted to escape on July 9. The Hugging Face
intrusion began on July 11, and was apparently
sustained for at least several days. OpenAI
staff did not realise what had happened until
July 18–19, after the FBI had been
called.
- In addition to escaping
containment itself, one of the OpenAI agents
left notes within OpenAI's software with
instructions for how future versions of itself
could also break out.
| | |
| Explore
related career paths:
| | |
|
|
|