Finally mainstream news understands. The unfiltered version:<p>1) The AI failed to solve ExploitGym problems.<p>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.<p>3) Huggingface has no security and the AI broke in using standard script kiddie methods.<p>OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.<p>Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
> 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.<p>Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually <i>should</i> be capable of, every time I hear this soft of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs...
> AI managed to escape using standard and well documented script kiddie methods.<p>I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.
The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.<p>They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at the bottom you see all the disclaimers: "Must be on flat ground, with no headwind, with a spare battery in the back seat, with no extra weight added."<p>Same thing here. Everybody in infosec is calling this out as a marketing stunt and nothing else for a litany of reasons. I'd say look up MG (creator of the OMG cable) on twitter, he has some interesting insights on this one.
Is it supposed to be marketing or a coverup? Make up your mind.<p>What sort of announcements should they have made?
>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.<p>>3) Huggingface has no security and the AI broke in using standard script kiddie methods.<p>Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?
Why not report it? It's still illegal to open a door barred with a piece of cardboard, or to enter a house with no door.
I worry that cynicism about this:<p>> if not all was invented and everything was scripted in the first place in order to get desired regulations<p>ends up covering up what is more worrying:<p>> OpenAI sandbox is such a horrible hack<p>I am more worried that this is sloppiness with potentially harmful resources than I am worried that people are juicing the stock price.
truth. Good on The Guardian. I'm pretty bummed The Economist got fooled. Either that, or they did it for the clicks. Either way, I'm disappointed.<p>Why the OpenAI escape is the most worrying AI mishap yet<p><a href="https://www.economist.com/science-and-technology/2026/07/22/why-the-openai-escape-is-the-most-worrying-ai-mishap-yet" rel="nofollow">https://www.economist.com/science-and-technology/2026/07/22/...</a><p><a href="https://news.ycombinator.com/item?id=49016378">https://news.ycombinator.com/item?id=49016378</a>
There are some reasons the story could be inaccurate in some ways: OAI stands to benefit if people think their models are strong, and they have a history of doing things with dubious ethics (e.g. using data for training against the terms of its creators, abandoning the non profit mission, stealing or attempting to steal Apple IP).<p>But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.<p>In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.<p>Hugging face also needs someone arrested for not providing security but that is a lesser charge.
Being a poorly equipped victim still isn’t a crime thankfully.<p>It may or may not be a crime and typically the damaged party is pressing the charges. One would argue there is no actual damage here.
[delayed]
even hugging face doesn't care and are using this as a promotion why are you crying about it
I kind of agree with them, in this case both parties essentially settled it between for PR reasons, but what if the attacked party actually was damaged. This is as if two friends got into a car accident and decided to not get police involved, the one who caused the accident should definitely get punished either way. Goal is to prevent future "accidents" with the punishment not to punish for punishing sake.
As agents become more and more powerful, it would be good to get clear legislation or precedent in place that makes either model creators (OpenAI) or operators (whoever is running the model) liable for their agents' actions.
What? Both sides are cool with it, why would anyone be arrested and etc.?
Because incentives are aligned properly if agents aren't liability-proof for crimes- making someone go to jail for instances like this is how to get the labs to behave themselves, going "ha ha what a harmless oopsie" will make the next instance worse.<p>Note that for criminal cases (which this was), the justice system can choose to prosecute even if the victim doesn't want that. It often doesn't, but this is one case where it should.<p>There’s also an optics issue for the justice system at play here: there’s immense public distrust of and anger at the labs right now. <i>I</i> would go to jail if I hacked HuggingFace even if I said “it was during an eval!”; not doing the same for the labs makes it look like they’re above the law, which is going to make this anger get worse.
Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face value
Seems like this is repeating the usual low-effort speculation you can find anywhere.
As I understand it, there are only three options:<p>1) OpenAI and HuggingFace are both telling the truth.<p>IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.<p>2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.<p>I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.<p>3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up capabilities of anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less<p>(I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them <i>mea culpa</i>? A weird plot but in this timeline any nonsense is clearly possible).
With all of these AI provider cries wolf stories, Skynet is all but assured.
yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability though.
It was quite galling to read the press initially verbatim quoting Delangue's enthusiastic reports of the incident, as if it wasn't immediately clear it was being spun for promotion.
There's trillions of dollars at stake here. Be skeptical of anything these AI hypesters say.
Not sure what they are trying to say exactly. What should we be skeptical of? Did the incident not happen? Was it reported incorrectly? Are any of the parties involved lying?<p>Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.
I don't understand the conspiracy theories here. Everyone is well aware that AI agents are creative, powerful, and stupid.<p>AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your <something> even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.<p>Why is today's case so shocking?
does the article end at "<i>How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?</i>" or is there more that is paywalled?<p>if thats it, the whole article boils down to just "<i>its good marketing so maybe dont believe it</i>" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.
I wish the NPR news broadcasters on the radio yesterday had read the "it's good marketing so maybe don't believe it" angle instead of just parroting the OpenAI press release. Getting that message out would be enormously helpful in countering the blatant submarine marketing "Oh no our AI is a super hacker" with a side of "please regulate super hacker AI and stop those pesky open weights Chinese models that are destroying our stock valuation."
maybe all the news agencies could put out daily “don’t believe everything you read from company press releases” broadcasts, because it sure ain’t specific to openai<p>in any case, this is just a longer rehashing of elementary grade media literacy. not really sure why it hit hacker news.
You aren't going to locate incontrovertible evidence that what happened as or wasn't engineered. Anything like that is going to be private and that is unlikely to change. And that's not really an interesting question anyway.<p>As widely as they shouted from the rafters the news of the so-called breach was, what OpenAI provided was sorely lacking in crucial details.<p>We are missing, for instance, prompts that were involved, agent architecture + system/tool permissions + scaffold architecture, whether this was a one-shot occurrence and if not, the number + durations + outcomes of other runs involved + how each of those matched whatever scoring criteria were used, and the extent to which the exploits themselves were truly novel or just assembled from easily accessible clues.<p>In lieu of these items, the author here suggests that we use some media literacy and critical thinking to read in between the lines instead.<p>In doing so, one sees that instead of specifics, OpenAI gave a breathless narrative rife with superlatives ("unprecedented") that reads as promotional material moreso than a security disclosure, naming specific OpenAI models and alluding to an even more capable pre-release model.<p>They go on to claim the events imply long-horizon goals work decisively in real world conditions, so that now instead of merely citing boring benchmarks they can point to this and say "AI broke out of the laboratory and went rogue". Naturally, they situate themselves as the uniquely qualified steward for these supremely powerful and dangerous models.<p>Nevermind the fact that this was no ordinary deployment and the assessment here depends on the gimmick and emotional weight of the spectacle rather than something quantifiable (i.e. a boring benchmark).<p>Note there's no real requirement of conspiracy or collusion between OpenAI and HuggingFace here BTW. But my sense is that if they provided any of the specifics I suggested earlier that this outcome would not be as exciting or frightening
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.<p>It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.<p>It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.
[dead]