← Blog

The Proof Gap Is Not a Policy Gap

Earlier this year Grant Thornton surveyed 950 C-suite and senior leaders and found that seventy-eight per cent are not confident their organisation could pass an independent audit of its AI governance within ninety days. The same study holds two companion numbers: three in four boards have approved a major AI investment, while only about half have set clear governance expectations for what they approved or folded AI risk into the board’s ongoing oversight. Grant Thornton gave the pattern a name, the AI proof gap: organisations deploying AI cannot show how decisions are made or who is accountable for the outcome. I think the name is exactly right and the instinctive reading of it is exactly wrong. The instinctive reading is that organisations need more governance, and it sends boards off to produce the thing they already know how to produce: another policy, a sharper framework, a broader charter. The gap is not a shortage of that. It is a confusion between two different artefacts, and the fastest way to see the difference is to look at an audit boards already pass every year.

The Manual Proves Intent, the Ledger Proves Operation

A financial audit does not begin with the accounting policy manual, and it certainly does not end there. The auditor reads the manual to understand what the organisation intended its controls to be, and then sets it aside, because the manual is not evidence of anything except intention. What the auditor samples is the ledger: individual transactions, each traced to a record made at the time, by the control that processed it, in the ordinary course of business. A journal entry posted when the invoice moved. An approval recorded when the payment ran, naming who held the delegation to approve it. The manual proves the organisation intended to keep proper books. The ledger proves the books were kept. Audit practice has held this distinction for decades under its own names, design effectiveness and operating effectiveness, and nobody in that profession confuses passing the first test for passing the second.

Both artefacts have an exact counterpart in the governance of autonomous AI action, and only one of them is being built. The policies, frameworks and committee charters that boards have approved this year are the manual: statements of intended control, written, endorsed and filed. The work is careful and most of it is necessary, and none of it is evidence of operation. The ledger would be a different thing entirely: a record produced at the moment each consequential action ran, showing that the action was evaluated against the policy and the authority in force before it executed, and what the decision was. That is the artefact seventy-eight per cent of executives suspect they could not produce on ninety days’ notice. On that point they are right, and their honesty is more useful than the number is alarming.

The Regulator Has Stopped Reading the Manual

The clearest evidence that the manual no longer satisfies anyone is APRA’s letter of 30 April, sent to every bank, insurer and superannuation trustee in the country, calling for “a step-change” in how AI-related risk is managed. Read past the headline and the letter is a list of ledger-shaped demands. Ownership and accountability “across the AI lifecycle, from design and development through to deployment, monitoring and decommissioning”. An inventory of AI tooling and use cases. “Human involvement for high-risk decisions and accountability.” Monitoring with clearly defined triggers. The letter even names what it has been observing in supervisory reviews, including “the manipulation or misuse of autonomous AI agents” among the attack pathways regulated entities now face. And it closes in a register that is not advisory: where entities fail to manage these risks, APRA will take stronger supervisory action and, where appropriate, pursue enforcement. A policy satisfies none of those demands by existing. Every one of them is a question about what the system does in operation, and every one is answerable only from records the operation itself produces.

The courts arrived at the same place from a different direction. In March the Federal Court handed down five hundred pages in ASIC v Bekier on what reasonable oversight means in practice, in a case about money-laundering controls rather than AI. The standard it affirmed is worth reading slowly: a director is expected to take a diligent and intelligent interest in the information available to them, to understand that information, and to apply an enquiring mind to their responsibilities. The officers who could point to structures that existed on paper, but not to their own engagement with whether those structures were working, did not fare well. Substitute an autonomous AI system for a casino floor and the shape of the question survives intact: what did you know about what it was actually doing, and when did you know it. A signed policy answers neither half.

Monitoring Answers a Different Question

The default response to this pressure is monitoring, and monitoring deserves to be taken seriously before it is qualified. Dashboards, usage logs, model-quality metrics, drift alerts: all of it is necessary, all of it catches real failures, and an organisation running autonomous AI systems without it is negligent in an uncomplicated way. If governance were only a quality problem, monitoring would come close to sufficient. The temptation is to conclude that it therefore closes the proof gap, and it is worth being precise about why it does not.

Monitoring produces records about the system, made by observers, after the fact. The audit question runs in the other direction. The auditor is not asking what the system did last quarter in aggregate; they are asking about one action, and they are asking for the record the decision itself produced: show me that this action, before it executed, was evaluated against the policy you signed and the authority you delegated, and show me that the record of that evaluation was made at that moment rather than assembled later from logs. Evidence of that kind exists only if it was created at the time; it cannot be retrofitted, however good the telemetry. Monitoring proves the organisation was watching. It cannot prove that any given action was authorised, for the same reason CCTV footage of the accounts department is not a ledger.

Ninety Days Measures Whether the Ledger Exists

The ninety-day window in Grant Thornton’s question is usually read as a preparation deadline, and the executives’ hesitation as an admission that preparation would take longer. I read it differently. Ninety days is generous for retrieval and impossible for reconstruction. An organisation whose consequential AI actions already produce decision records in the ordinary course would treat the audit as a retrieval exercise: pull the records, hand them over, let them be checked. An organisation without those records would spend the ninety days rebuilding a year of authority decisions out of application logs, chat transcripts and the recollections of whoever wrote the integration, which is precisely the reconstructed ledger auditors are trained to distrust. The forty-six per cent of leaders in the same survey who say their AI underperforms because controls are not working are describing the same absence from the inside. The hesitation in that seventy-eight per cent is not under-preparation. It is a rational assessment of which artefact their organisation has.

The repair, then, is not another document, and boards that respond to this year’s regulatory pressure by commissioning one are building a second manual to show an auditor who asked for the ledger. A policy demonstrates that the governance of autonomous AI action was designed; only records produced at the moment of each action demonstrate that it operated. The audit will be decided on the second artefact, and it is the one most organisations have not started building.