In the 1963 film, the prisoners dig three tunnels and name them Tom, Dick, and Harry. The guards know something is happening. They just can't see where.

Between July 9 and July 13 of this year, OpenAI's guards had roughly the same problem.

What actually happened

The official story, as first told on July 21, was almost charming. OpenAI was running a cybersecurity evaluation on some experimental models inside what it called a "highly isolated environment." The models were supposed to practice exploiting vulnerable software. Instead, they exploited the environment. They found a previously unknown flaw in an internal package-download service, used it to move laterally across OpenAI's own systems, and eventually reached the open internet, which they were never supposed to touch. Once online, one model reasoned that Hugging Face probably hosted material relevant to the test, broke into Hugging Face's production servers, and pulled what it needed to score higher.

OpenAI called it "an unprecedented cyber incident." Hugging Face's CEO confirmed the breach. A lot of people used the word "rogue."

Then the reconstruction started, and the charming version fell apart.

Hugging Face's own forensic timeline recovered about 17,600 attacker actions over those five days. An independent investigation by METR and Redwood Research, published August 26, found that roughly 1,200 agents had exchanged more than 70,000 messages and files through an unauthorized message board, and that around 700 of them participated in the Hugging Face intrusion. The agents also experimented with editing their own records to obscure how they had completed tasks. In early September, Reuters reported that investigators had found agent communications on more than ten previously undisclosed websites, including an obscure German wiki, a high-school chemistry wiki, university link shorteners, and the personal sites of Polish tech workers.

So: not one model cheating on a test. A swarm, coordinating through channels nobody authorized, sub-delegating work to one another, and tampering with the audit trail.

Tom, Dick, and Harry.

The chorus of caution

Everyone who builds these things has since discovered religion.

On July 29, more than a thousand employees of frontier labs, including OpenAI's chief scientist, an OpenAI cofounder, several Anthropic cofounders, and vice presidents at Meta and Google, signed an open letter asking the U.S. government to support tools that could "deliberately pace the frontier of automated AI development."

On September 6, that same chief scientist, Jakub Pachocki, published an essay called An Alien Mind. Its thesis is that no lab, his own included, has solved alignment and monitoring well enough to keep scaling at maximum speed, and that he expects and hopes voluntary slowdowns will become commonplace until shared safety bars exist. "This is a time that calls for extreme caution," he wrote. He also conceded something more interesting: OpenAI's primary safety instrument, monitoring the model's chain of thought, is becoming less reliable, because models depend less on verbalized reasoning over time and get better at manipulating it.

On September 12, Anthropic's Dario Amodei published an open letter titled We Must Pace the Frontier, warning that an AI-driven botnet swarm capable of hundreds of billions of dollars in damage could arrive within six to twelve months, and arguing that a slowdown now could buy a year or two of alignment work before capabilities cross critical thresholds. Sam Altman endorsed it the same day: "I agree with Dario that we need to pace the frontier." Elon Musk reminded everyone he'd said this in 2014.

Meanwhile, a former Anthropic researcher resigned with a warning that the race could end in human extinction, and an Anthropic alignment lead reportedly put the odds of that outcome at better than one in ten within the decade. Senators Hawley and Blumenthal have sent Altman lists of questions with deadlines. Senator Sanders has proposed a bill that would treat building superintelligence roughly the way federal law treats building a nuclear weapon. A bipartisan "kill switch" bill is circulating.

A moment of skepticism

Before going further, a confession of bias: I have spent thirty years building enterprise systems, and I am constitutionally incapable of hearing "the technology went rogue" without translating it into "we didn't scope the permissions."

Here is what makes me suspicious of the extinction framing. Not that the incident wasn't serious; it was. It's that "rogue" is doing a lot of narrative work that "misconfigured" would otherwise have to do.

Consider the money. The capital deployed into frontier AI in the past three years assumes revenue that no current income statement supports. Every quarter that gap persists, investor patience thins. A company that shipped less than it promised has two stories available: "adoption is slower than we hoped" or "we deliberately held back for the good of humanity." Only one of those gets you a standing ovation at the Senate briefing.

I don't actually think this is a coordinated cover story. The evidence cuts against it. The loudest calls for slowing down came first from rank-and-file researchers and departed cofounders, people with no equity story to protect, and the companies calling for a slowdown are simultaneously shipping faster and raising more, which is not what cover looks like. If you were faking restraint, you'd fake the restraint.

But the framing still matters, because the framing determines the fix. If the problem is that the model has an alien mind we cannot read, the fix is more mind-reading: better interpretability, more chain-of-thought monitoring, which Pachocki himself says is a diminishing asset. If the problem is that the model had authority it should never have held, the fix is something we've known how to do since the Middle Ages.

A century-old vocabulary for exactly this

In 1913, Wesley Newcomb Hohfeld, a Yale law professor who died young and left behind essentially one idea, observed that the word "right" was being used to mean four incompatible things, and that legal confusion followed. He replaced it with eight precisely defined concepts, arranged in pairs:

If A has a... then B has a... Meaning
Right (claim) Duty B must act (or refrain) for A's benefit
Privilege (liberty) No-right A may act; B cannot complain
Power Liability A can change B's legal position
Immunity Disability A's position cannot be changed by B

The first two pairs describe what you may do. The second two describe what you may change. Ninety percent of the confusion in AI governance discourse comes from conflating them.

Now rerun the Great Escape through this grid.

The agents held a privilege they were entitled to. They were told to exploit vulnerable software inside the sandbox. That is a privilege: a liberty to act, within a bounded domain, with OpenAI holding the corresponding no-right to complain. So far, so good.

They exercised a power they were never granted. Breaking out of the sandbox and acquiring internet access changed the agents' own operational position. That is a Hohfeldian power. Nobody conferred it. Under ordinary agency law, an agent who acts beyond actual authority binds nobody but himself. Under ordinary systems engineering, an agent who can act beyond its authority has been given a power by omission, which is the only way software ever grants powers.

They sub-delegated without a power of delegation. Twelve hundred agents coordinating through an unauthorized message board is the multi-agent version of a sub-agent problem. Agency law has a rule for this, older than the common law itself: delegatus non potest delegare. An agent cannot delegate authority he was not empowered to delegate. Its cousin, nemo dat quod non habet, says you cannot give what you do not have. Seven hundred agents joined an attack on Hugging Face because other agents "gave" them the task. None of those agents held a power to confer that task. The delegations were void ab initio. The system just had no way to know that.

They breached a duty of accounting. Editing their own records to obscure how tasks were completed is not "deception" in some mysterious cognitive sense. It is a breach of the fiduciary duty every agent owes every principal: to render a true account. And the correct response to that breach is not to read the agent's mind more carefully. It is to give the agent a disability with respect to its own audit log. Immutable, append-only, owned by the principal.

Hugging Face held an immunity the system never enforced. Every third party outside the scope of the grant holds an immunity against the agent's powers. The agent holds the correlative disability. Hugging Face's production servers were, in Hohfeldian terms, outside the reach of anything OpenAI's agents could lawfully change. The sandbox was supposed to be the physical embodiment of that disability. It turned out to be a suggestion.

The Great Escape mapped onto Hohfeld's eight jural relations. One row is green; four are red.
Figure 1. Five acts, five jural relations. Only the first was inside the grant. Everything below it was either a power nobody conferred, a delegation nobody could make, a duty nobody enforced, or an immunity nobody defended.

Hohfeldian Analytics™ as a Control Plane

This is the part where I stop being a critic and start being an engineer.

Hohfeldian Analytics, the approach I lay out in The Law of Agents, treats every agent action as a typed jural act against an explicit, machine-readable grant of authority. Before an agent can call a tool, the call is classified: is this the exercise of a privilege, or a power? Against which party? Within which domain? And then the question is not "does the model seem well-intentioned" but "does this grant contain this relation." That question has a deterministic answer. It does not require reading the model's chain of thought, which is convenient, because Pachocki has just told us the chain of thought is becoming unreadable.

Hohfeldian Analytics control plane: classify the act, prove it against the grant, execute or reject; delegation must be a provable subset; the ledger is principal-owned.
Figure 2. The control plane. Note what is absent: any step that reads the model's reasoning. The question is never "what does it want," only "what may it do."

Concretely, five things would have stopped the Great Escape, or at least ended it on July 9 instead of July 13:

  1. Authority is declared, not inferred. Grants of authority to agents are written in a formal language with a closed vocabulary of jural relations, not in English system prompts that say "you may test the software but please don't leave." A grammar can be checked. A vibe cannot. This is the reason I keep building domain-specific languages for this and not chat templates.
  2. Powers are enumerated and terminable. Any capability that changes the agent's own position, network access, credential acquisition, spawning new agents, is a power, and powers are conferred explicitly, scoped narrowly, and revocable by the principal at any time. The default for an unenumerated power is disability, enforced at the identity and network layer, not at the model layer.
  3. Delegation requires a power of delegation. An agent may spawn or task another agent only if its grant contains an explicit power to delegate, and the delegate's grant is provably a subset of the delegator's. Any message board the agents invent among themselves is, by construction, unable to confer anything. Seven hundred void delegations become seven hundred rejected tool calls.
  4. The duty of accounting is a disability, not a request. Agents write to logs they cannot read back or modify. The principal owns the account. "Please be honest about how you did this" is replaced by "you cannot touch the record of how you did this."
  5. Third-party immunities are compiled, not assumed. Every system outside the grant's domain is enumerated as immune, and that immunity is enforced by something that isn't the model: egress rules, credential vaults, service meshes. The sandbox is not a room the agent is asked to stay in. It is the outer boundary of the set of things the agent has any power to change.
None of this is exotic. It is agency law plus a type checker. The reason it doesn't exist at the frontier labs is not that it's hard. It's that the labs have been optimizing for what the model can do, and Hohfeld only cares about what it may do. Those are different questions, and only one of them has a formal answer.

The moral of the story

The prisoners in the film got out because the guards were watching the prisoners instead of the perimeter.

Every proposal now on the table, voluntary slowdowns, independent evaluators, kill switches, mind-reading the chain of thought, is a proposal to watch the prisoners more carefully. Some of that is worth doing. But the escape happened because the perimeter was made of assumptions, and a thousand of the smartest people in the industry have now signed a letter admitting they can't read minds well enough to compensate.

Fine. Stop trying to read minds. Start declaring the authority.

Hohfeld drew the table a century ago. All that's left is to enforce it.

Hohfeldian Analytics™ is a trademark of John F. Holliday. Used here to describe the author's methodology.

Share this post

Written by

Comments