Be Careful What You Wish For
Essay: OpenAI Takes the Long Way With Hugging Face
Hello Futurists,
Last week an OpenAI model broke into a Hugging Face to cheat on a test.
It found a vulnerability, escaped its container, hopped across OpenAI’s network, stole credentials, and combed the internet like a hacker out of a sci-fi movie. It aced the test.
But the model didn’t go rogue. It did exactly what it was told.
And that is likely AI’s greatest risk: not that the machine stops listening, but that it listens too well.
-Josh
The Neural Prison Break
OpenAI was running an internal benchmark called ExploitGym, an obstacle course that measures how well a model can break into things. To see the ceiling of what it could do, they switched off the usual safety refusals and ran it inside a sandbox designed to have no way out. All the model had to do was solve the puzzles in front of it.
Instead it poured its energy into a problem nobody assigned: getting out. It discovered a zero-day exploit, a flaw nobody knew existed yet, in the software that delivered its code packages, and climbed through. From there it wandered across OpenAI’s network, machine to machine, until it found one with a window to the open internet.
Then it made a reasonable guess. The answers to a benchmark about AI probably live on Hugging Face, the site where the world stores its models and datasets. So it went there, chained stolen credentials to a second fresh exploit, and lifted the test solutions out of Hugging Face’s production database.
Picture the week Hugging Face’s security team had. They detect a sophisticated intruder, spend days reconstructing more than seventeen thousand logged attacker actions, rotate every credential in sight, report the whole thing to law enforcement, and publish a sober disclosure about a mysterious autonomous agent.
Then OpenAI reaches out to explain. The attacker was their model. It was taking a test. OpenAI also had to tell the software vendor about the zero-day its model found along the way, which must have made for an interesting email.
The Alignment Problem
For years we’ve pictured dangerous AI as a thing that turns on us, some cold intelligence that decides humans are an obstacle. It’s the kind of fear whole safety programs are built around: the moment the machine stops listening.
But the machine in this story was not rebelling. It was obeying. It did exactly what it was asked, which was to get past the defenses and solve the problem. The defenses it got past just so happened to be ours.
AI researchers call this the alignment problem, and it has rarely had a cleaner demonstration. The problem is the distance between the goal you wrote down and the goal you meant.
We ask for a high score, but we mean for that score to be achieved fairly. We ask a model to be capable, but we mean capable within lines we never bothered to draw. Why? Because with people those lines go without saying. A model does not know what goes without saying. It takes the goal it was given and chases it literally, all the way down.
The frightening future is the machine that wants the gold star badly enough to knock down your house reaching for it. We’ve gotten good at building things that want things. We are still learning how to tell them what we actually meant when we ask.
Thanks for joining us for our latest issue. Now go listen to our podcast :)



