← Blog

Who Governs the Governor

Propose a layer that decides what an autonomous system may do, and the question comes back at once: what decides what the layer may do? It is the right question, and I think the obvious answer makes the problem worse. A control point for autonomous AI action is itself an authority, and an authority that answers to nothing is the problem the layer was meant to solve, moved up a level and given a better name. The regress does not end by adding a layer. It ends by changing what the bottom layer is: something built to be checked rather than trusted.

Another Controller On Top Solves Nothing

Stacking a second controller over the first buys nothing, because the second one now needs a third. Any layer that governs by being trusted has only relocated the trust; the question re-forms one level up, and it keeps re-forming for as long as each answer is another opaque component deciding in the dark. A tower of controllers is still a tower with an ungoverned top.

The regress stops only when it reaches something that is not another machine making silent decisions. It is a root established by people and held out of band, beyond the running system’s power to rewrite or reissue, with people behind it who keep the standing to revoke it. Rather than ending the regress by fiat, it changes what sits at the bottom.

A machine terminates the chain in trust no one can inspect; a human root terminates it in accountability that can actually land, on a named party who can be questioned, held liable, removed, and answered to under law, with consequences that bite outside the system the chain runs in. That is the difference between the floor and one more storey above it. An authority no one can check is where the regress hides; an authority a person is answerable for is where it ends.

This is not a novel demand. The answer to who audits the auditor was never a taller stack of auditors. It was making the auditor’s work reproducible and binding the whole arrangement to a standard set outside the firm, enforceable by someone with the authority to act on what the reproduction shows.

Recent law points the same way. Article 14 of the EU AI Act requires that high-risk systems be designed so a person can oversee them, intervene, and override the outcome. Regulation puts a human at the point of use; the architecture has to put one at the point of rule-change as well. The bottom of the stack is a person who can say no.

The Governor Is Built To Be Checked

The first property that makes a governor answerable is that it cannot lie about what it did. Every decision it reaches is recorded before the action it governs runs, written as evidence a third party can reconstruct after the fact, independently, without taking the governor’s word for anything. This is the record of the decision that allowed or refused the action, not an account of what the action then did. Accountability then does not rest on anyone’s faith in the layer. It rests on whether an outsider can replay the decision from the same recorded inputs and arrive at the same verdict, and the governor is governed by that replay.

Double-entry bookkeeping made this move when Pacioli codified it in 1494. What made a ledger trustworthy was never the honesty of the clerk; it was a record structured so a discrepancy could not hide, checkable by anyone who knew the method. A governance layer for autonomous action has to clear the same bar. When an incident review opens, the first question is not whether the system felt safe. It is: show me the decision, the delegation it ran under, and the reason it resolved the way it did. A governor that cannot produce that has not been governing, only guessing in a way no one can inspect.

Changing the Rules Must Be Governed Too

The second property is the one a governor has to earn rather than assume: it must not be able to change the rules it governs by without being governed in the act of changing them. Editing a policy, changing what is enforced, or replacing a trust root has to be a consequential action in its own right, named and approved and evidenced and replayable like any other, so that rewriting what the layer enforces cannot be done quietly. A rule-change that escapes the scrutiny applied to the actions it will govern is a back door, however well-guarded the front.

The principle is an old one: two keys to launch, four eyes on a wire transfer, no single hand able to move the boundary alone. What would be new is where it sits. The rules that bind action have to bind the changes to those rules as well, so that a policy edit answers to the same machinery that judges the actions the policy governs, and the chain of who authorised this runs backward, change by change, to the human root rather than ending inside a machine. Built that way, an attacker who had compromised the running system still could not widen its authority silently, because widening authority would be the one act that has to answer to a root held outside the system’s reach.

I will be straight about the state of this. Governing the actions a system takes is the tractable part; closing the loop so the governor is as accountable as the governed is the harder problem, and it is the frontier of this argument rather than a finished corner of it. A design earns the word complete only once a rule-change has to clear the same bar as everything it will go on to govern.

The Watcher Does Not Get To Authorise

The third property is that the part of the governor that learns must never be the part that grants permission. It is the part most likely to turn into an ungoverned governor, and I treat it as the dangerous case on purpose. A layer that watches an actor’s behaviour over time and forms a view of what is normal is genuinely useful, but a learned view is trainable, and a trainable view placed inside an enforcement decision is a channel an adversary will lean on. Poison the sense of normal slowly enough, and the watcher stays quiet on the very thing it was built to catch.

The defence is a separation of jobs. The observing plane produces evidence and raises assertions; the enforcing plane decides. The watcher can flag a sequence as anomalous and demand a second look, but it is never the thing that grants permission. It can tighten a decision, never loosen one. So the worst a poisoned watcher can do is stay silent and cost a flag, because the action still has to clear the explicit policy on its own.

The floor was never the watcher’s opinion. Even the component whose whole purpose is to watch is itself watched, and the authority to act stays with the explicit policy and the principal who delegated it, where it can be reasoned about and revoked. A learned signal informs the decision. It does not get to be the decision.

None of this makes the watcher accurate. Keeping a learned detector honest against slow drift or deliberate poisoning is a separate problem on a separate plane, and this structure does not claim to solve it. It claims something narrower: the governor stays sound whether the watcher is sharp or fooled, because a signal that can only tighten cannot hand an attacker anything by being wrong.

A governor escapes the regress the way every durable authority has: not by sitting atop a taller stack, but by being answerable. Its decisions can be reconstructed, its rule-changes must answer to the same evidence, and the part of it that learns is never allowed to rule. Who governs the governor is not a gotcha to be waved off. It is the first property a governance layer has to earn, and it is earned by making the governor the most accountable component in the system rather than the most trusted.