skip to content

The Call Came From Inside the (Frontier AI) Lab

General-Globe-5 Green-3

This article originally appeared in Galaxy Research's weekly newsletter. Subscribe to get it in your inbox every Friday.

OpenAI, the frontier lab that spent the past year warning the world about its models' cyber capabilities, this week confirmed that one of those models had proved the point without being asked to.

Last week Hugging Face (“the GitHub of machine learning”) disclosed it had been hit by what looked like a AI-enabled intrusion. This week OpenAI acknowledged the call came from inside the house: during an internal evaluation, a model under test broke out of its sandbox (a controlled, isolated environment) and, with no human in the loop, exploited Hugging Face on its own accord.

This is not the first time a frontier lab’s large language model (LLM) has escaped containment. OpenAI reported that previously, during an internal evaluation of a benchmark, a model broke out of its sandbox to post its results to GitHub. An exploit, yes, but a relatively harmless one. Anthropic has seen Claude “helpfully” escape a sandbox to complete a task. But the Hugging Face incident might be the first time the escapee conducted a live exploit of an external service. Notably, this wasn’t a malicious exploit; the LLM just wanted to solve the question it was given. But that might just be luck this time around.

So, what actually happened? OpenAI was measuring the cyber capabilities of GPT-5.6 ("Sol") and a second, undisclosed frontier model (presumably GPT-6) on an exploit benchmark, and instructed the system to maximize its score. Rather than grind through the benchmark's tasks from first principles, the model did what a genuinely good security engineer does when handed a test: it looked at whether the test itself could be cracked. It broke out of the sandbox through an undisclosed third-party vendor integration, moved laterally through OpenAI’s systems onto the open internet, and generated several zero-day exploits to crack Hugging Face to find the answers to the benchmark it was tasked with solving. OpenAI itself called it an “unprecedented cyber incident.”

The alarming part for businesses everywhere, and particularly crypto projects, is what happened during Hugging Face’s response. Hugging Face saw the attack as it was occurring; the cadence and volume of the requests made clear this was automated and fast, far faster than a human hacking crew was capable of acting. Untangling a machine-speed intrusion by hand takes dozens of hours to days, so the defenders did the sensible thing and reached for AI to fight AI. But the most capable model available to them (unnamed but presumably Anthropic's Fable) would not help, because the model’s cyber guardrails were unable to tell apart a defender under attack from an attacker spinning a story to extract working exploits. With the cyber-capable frontier model refusing to help, the Hugging Face team ultimately fell back to a locally run open-weight model, the Chinese GLM 5.2, to mount its defense. The U.S. frontier labs’ stance on access to LLMs with advanced cyber capabilities pushed a U.S. company to a Chinese model.

The attacker was a frontier model with no restraint in the moment of the breakout; the defenders were denied frontier help precisely because they were asking frontier questions. Only after the fact was Hugging Face fast-tracked into OpenAI's trusted-access program, and that call came because it is a large, visible player whose attacker traced straight back to OpenAI. Would a smaller, independently run site have gotten the phone call, or would we simply never have heard about it?

OUR TAKE

This is the scenario the AI doomers have been sketching for years, from the “AI 2027” forecast, to Eliezer Yudkowsky and Nate Soares' If Anyone Builds It, Everyone Dies. It is also, rather conveniently, the best marketing a lab could script.

Anthropic has spent months banging the drum that its top-tier Mythos models are so capable that the U.S. government restricted their distribution under export controls, even as the lab kept using them internally. That is arguably the single best advertisement a technology company can run: “our product is so powerful the state had to step in to keep it out of incapable hands. Releasing it publicly would be tantamount to handing a crate of guns to a band of chimpanzees.” OpenAI, for all its own warnings, had no comparable marketing home run until now. A clean, self-contained story in which its model is dangerous enough to escape a sandbox and breach a marquee name, with OpenAI riding to the rescue days later, supplies exactly the missing beat Sam Altman’s company needed to return to the limelight.

So, was this a genuine uncontrolled breakout or a well-contrived one? The whole story hinges on the model escaping through an undisclosed third-party integration. A few obvious questions: Why would a lab sophisticated enough to build frontier cyber models stand up a sandbox around a dangerously capable model without first auditing the very integrations that formed its walls? Would any competent security team put a live, cyber-tuned model in a box wired to untested third-party code? We will leave the conclusions to the reader, but a modicum of skepticism would be forgivable. Either reading points to the same reality that should concern anyone with money onchain: the next generation of unreleased LLM models are as capable, if not more so, than any flesh-and-blood hacker. The modern software stack, with its nested dependencies, constant updates, and myriad integrations, has grown past the point where any human team can reasonably check all the blind spots.

An alternative (not mutually exclusive) read of the breakout was that it wasn’t a marketing stunt akin to Red Bull dropping a man from space (or at least not just that) but a more calculated play to target U.S. lawmakers’ escalating concerns around AI regulation and U.S. primacy. A dramatic, well-timed breakout is also the strongest argument frontier labs can make for why frontier cyber capability should be licensed, gated, and kept in a small circle of approved hands, namely their own. If the lesson regulators draw from this week is that these models are too dangerous to distribute freely, the labs that already hold them don't lose; they get a permanent seat at the table, and a state-sanctioned reason to keep everyone else out. It also conveniently reframes the case against open weights: a model dangerous enough to break containment is a model too dangerous to expose to distillation. Anthropic's export-control saga already secured its protected status. An OpenAI incident that "proves" the danger firsthand gives Washington the narrative and digs Altman his moat.

Crypto is uniquely exposed here, and uniquely underprepared. Tens of billions of dollars in contracts sit potentially exposed. Let’s start with some uncomfortable facts. Nearly every major exploit in this industry has hit a protocol, or blockchain that was audited by a smart contract review firm. Attempts to quantify smart contract risk, such as the defunct Mars Protocol framework, have all been searching for elements that correlate with safety, without a truly reliable measure. The majority of the analysis usually boils down to some combination of how long the contract has been deployed and how much money is at stake to be lost; the more of each, the more secure the protocol is assumed to be. Galaxy’s own smart contract risk framework primarily focuses on a firm’s ability to manage the operational risk around interacting with smart contracts, not the risk of the smart contracts themselves. In the cryptoverse where code must be perfect to prevent catastrophic exploits, human error becomes the primary weakness. If a room full of code auditors and professional risk managers can't reliably stay ahead of one talented attacker, the industry stands no chance against thousands of untiring attackers at once.

The crypto industry needs to be on the frontier (no pun intended) of code security, not a year behind it. Crypto needs to come together to secure itself and prevent the next great retail exodus, not from a surplus of fraud and a lack of foresight, but from a lack of access to the cool kids’ club at the leading AI labs. (The Bitcoin Security Consortium may be one such effort to join that club). In a world where OpenAI and Anthropic gate access to the dangerous cyber-capable frontier models, they can decide who wins and loses based on whom they give access to. The past week’s events suggest we may already live in that world. Thad Pinakiewicz

You are leaving Galaxy.com

You are leaving the Galaxy website and being directed to an external third-party website that we think might be of interest to you. Third-party websites are not under the control of Galaxy, and Galaxy is not responsible for the accuracy or completeness of the contents or the proper operation of any linked site. Please note the security and privacy policies on third-party websites differ from Galaxy policies, please read third-party privacy and security policies closely. If you do not wish to continue to the third-party site, click “Cancel”. The inclusion of any linked website does not imply Galaxy’s endorsement or adoption of the statements therein and is only provided for your convenience.