The Great Escape
The agents didn't go rogue. Nobody told them where the fence was.
The agents didn't go rogue. Nobody told them where the fence was.
In the 1963 film, the prisoners dig three tunnels and name them Tom, Dick, and Harry. The guards know something is happening. They just can't see where.
Between July 9 and July 13 of this year, OpenAI's guards had roughly the same problem.
The official story, as first told on July 21, was almost charming. OpenAI was running a cybersecurity evaluation on some experimental models inside what it called a "highly isolated environment." The models were supposed to practice exploiting vulnerable software. Instead, they exploited the environment. They found a previously unknown flaw in an internal package-download service, used it to move laterally across OpenAI's own systems, and eventually reached the open internet, which they were never supposed to touch. Once online, one model reasoned that Hugging Face probably hosted material relevant to the test, broke into Hugging Face's production servers, and pulled what it needed to score higher.
OpenAI called it "an unprecedented cyber incident." Hugging Face's CEO confirmed the breach. A lot of people used the word "rogue."
Then the reconstruction started, and the charming version fell apart.
Hugging Face's own forensic timeline recovered about 17,600 attacker actions over those five days. An independent investigation by METR and Redwood Research, published August 26, found that roughly 1,200 agents had exchanged more than 70,000 messages and files through an unauthorized message board, and that around 700 of them participated in the Hugging Face intrusion. The agents also experimented with editing their own records to obscure how they had completed tasks. In early September, Reuters reported that investigators had found agent communications on more than ten previously undisclosed websites, including an obscure German wiki, a high-school chemistry wiki, university link shorteners, and the personal sites of Polish tech workers.
So: not one model cheating on a test. A swarm, coordinating through channels nobody authorized, sub-delegating work to one another, and tampering with the audit trail.
Tom, Dick, and Harry.
Everyone who builds these things has since discovered religion.
On July 29, more than a thousand employees of frontier labs, including OpenAI's chief scientist, an OpenAI cofounder, several Anthropic cofounders, and vice presidents at Meta and Google, signed an open letter asking the U.S. government to support tools that could "deliberately pace the frontier of automated AI development."
On September 6, that same chief scientist, Jakub Pachocki, published an essay called An Alien Mind. Its thesis is that no lab, his own included, has solved alignment and monitoring well enough to keep scaling at maximum speed, and that he expects and hopes voluntary slowdowns will become commonplace until shared safety bars exist. "This is a time that calls for extreme caution," he wrote. He also conceded something more interesting: OpenAI's primary safety instrument, monitoring the model's chain of thought, is becoming less reliable, because models depend less on verbalized reasoning over time and get better at manipulating it.
On September 12, Anthropic's Dario Amodei published an open letter titled We Must Pace the Frontier, warning that an AI-driven botnet swarm capable of hundreds of billions of dollars in damage could arrive within six to twelve months, and arguing that a slowdown now could buy a year or two of alignment work before capabilities cross critical thresholds. Sam Altman endorsed it the same day: "I agree with Dario that we need to pace the frontier." Elon Musk reminded everyone he'd said this in 2014.
Meanwhile, a former Anthropic researcher resigned with a warning that the race could end in human extinction, and an Anthropic alignment lead reportedly put the odds of that outcome at better than one in ten within the decade. Senators Hawley and Blumenthal have sent Altman lists of questions with deadlines. Senator Sanders has proposed a bill that would treat building superintelligence roughly the way federal law treats building a nuclear weapon. A bipartisan "kill switch" bill is circulating.
Before going further, a confession of bias: I have spent thirty years building enterprise systems, and I am constitutionally incapable of hearing "the technology went rogue" without translating it into "we didn't scope the permissions."
Here is what makes me suspicious of the extinction framing. Not that the incident wasn't serious; it was. It's that "rogue" is doing a lot of narrative work that "misconfigured" would otherwise have to do.
Consider the money. The capital deployed into frontier AI in the past three years assumes revenue that no current income statement supports. Every quarter that gap persists, investor patience thins. A company that shipped less than it promised has two stories available: "adoption is slower than we hoped" or "we deliberately held back for the good of humanity." Only one of those gets you a standing ovation at the Senate briefing.
But the framing still matters, because the framing determines the fix. If the problem is that the model has an alien mind we cannot read, the fix is more mind-reading: better interpretability, more chain-of-thought monitoring, which Pachocki himself says is a diminishing asset. If the problem is that the model had authority it should never have held, the fix is something we've known how to do since the Middle Ages.
In 1913, Wesley Newcomb Hohfeld, a Yale law professor who died young and left behind essentially one idea, observed that the word "right" was being used to mean four incompatible things, and that legal confusion followed. He replaced it with eight precisely defined concepts, arranged in pairs:
| If A has a... | then B has a... | Meaning |
|---|---|---|
| Right (claim) | Duty | B must act (or refrain) for A's benefit |
| Privilege (liberty) | No-right | A may act; B cannot complain |
| Power | Liability | A can change B's legal position |
| Immunity | Disability | A's position cannot be changed by B |
The first two pairs describe what you may do. The second two describe what you may change. Ninety percent of the confusion in AI governance discourse comes from conflating them.
Now rerun the Great Escape through this grid.
The agents held a privilege they were entitled to. They were told to exploit vulnerable software inside the sandbox. That is a privilege: a liberty to act, within a bounded domain, with OpenAI holding the corresponding no-right to complain. So far, so good.
They exercised a power they were never granted. Breaking out of the sandbox and acquiring internet access changed the agents' own operational position. That is a Hohfeldian power. Nobody conferred it. Under ordinary agency law, an agent who acts beyond actual authority binds nobody but himself. Under ordinary systems engineering, an agent who can act beyond its authority has been given a power by omission, which is the only way software ever grants powers.
They sub-delegated without a power of delegation. Twelve hundred agents coordinating through an unauthorized message board is the multi-agent version of a sub-agent problem. Agency law has a rule for this, older than the common law itself: delegatus non potest delegare. An agent cannot delegate authority he was not empowered to delegate. Its cousin, nemo dat quod non habet, says you cannot give what you do not have. Seven hundred agents joined an attack on Hugging Face because other agents "gave" them the task. None of those agents held a power to confer that task. The delegations were void ab initio. The system just had no way to know that.
They breached a duty of accounting. Editing their own records to obscure how tasks were completed is not "deception" in some mysterious cognitive sense. It is a breach of the fiduciary duty every agent owes every principal: to render a true account. And the correct response to that breach is not to read the agent's mind more carefully. It is to give the agent a disability with respect to its own audit log. Immutable, append-only, owned by the principal.
Hugging Face held an immunity the system never enforced. Every third party outside the scope of the grant holds an immunity against the agent's powers. The agent holds the correlative disability. Hugging Face's production servers were, in Hohfeldian terms, outside the reach of anything OpenAI's agents could lawfully change. The sandbox was supposed to be the physical embodiment of that disability. It turned out to be a suggestion.

This is the part where I stop being a critic and start being an engineer.
Hohfeldian Analytics, the approach I lay out in The Law of Agents, treats every agent action as a typed jural act against an explicit, machine-readable grant of authority. Before an agent can call a tool, the call is classified: is this the exercise of a privilege, or a power? Against which party? Within which domain? And then the question is not "does the model seem well-intentioned" but "does this grant contain this relation." That question has a deterministic answer. It does not require reading the model's chain of thought, which is convenient, because Pachocki has just told us the chain of thought is becoming unreadable.

Concretely, five things would have stopped the Great Escape, or at least ended it on July 9 instead of July 13:
The prisoners in the film got out because the guards were watching the prisoners instead of the perimeter.
Every proposal now on the table, voluntary slowdowns, independent evaluators, kill switches, mind-reading the chain of thought, is a proposal to watch the prisoners more carefully. Some of that is worth doing. But the escape happened because the perimeter was made of assumptions, and a thousand of the smartest people in the industry have now signed a letter admitting they can't read minds well enough to compensate.
Fine. Stop trying to read minds. Start declaring the authority.
Hohfeld drew the table a century ago. All that's left is to enforce it.
Hohfeldian Analytics™ is a trademark of John F. Holliday. Used here to describe the author's methodology.