How Do You Know Your AI Governance Lead Is Not Just a Rubber Stamp?
Every regulated firm is now trying to hire the same person: someone who can make yes safe. Someone who gets AI into production without the institution regretting it. It is the right hire. But almost no one can answer the question a board should ask next: how would you know if that person is actually doing the job, or quietly rubber stamping?
This is the harder half of the problem, and it is where most firms will get it wrong. I have spent years on the delivery side of this, getting AI into production past the people whose job is to sign or to stop it. Here is what actually tells you whether that role is working.
The stakes are not abstract. Under the EU AI Act, and the model risk regimes that already govern banks, accountability for a deployed system attaches to a named person, not a platform and not a committee. If that person is a rubber stamp, the assurance the board is relying on is fiction, and the board is the one exposed when it fails.
The two obvious scorecards both fail
Start with how you would measure the role, because the intuitive answers are traps.
Measure the person on deployment speed, on how fast they get models live, and governance quietly dilutes under delivery pressure. Make yes safe becomes find a way to yes. A comment on my last piece named that failure mode, and named it more sharply than anything in the original. The person keeps shipping, the reviews get thinner, and no one notices until something that should not have gone out already has.
Measure them the other way, on control compliance and clean audits, and the role drifts back into assurance. It becomes the approval layer you were trying to escape, a function that is safe, slow, and accountable for nothing that ships. You have rebuilt the bottleneck and given it a new title.
The deeper problem is that metrics get optimized toward whoever writes them. Hand the scorecard to delivery and it rewards speed. Hand it to risk and it rewards caution. No single scorecard, owned by one side, survives contact with the incentive it creates.
The answer is an architecture, not a scorecard
So stop looking for the perfect metric. This role is governed by how you structure it, not by what you measure.
Dual authorship. Delivery owns the budget and the headcount, so the incentive points at shipping. Risk countersigns the objectives, not as a veto, but as a signature on what good looks like. The tension between speed and safety has to be written into the objectives themselves, not resolved in advance in favor of one side. If only one function authors the goal, the goal will serve that function.
The ungameable pair. Track two numbers together: control exceptions caught before release, and incidents found after. They move in opposite directions under manipulation. Dilute governance to ship faster and the incidents after release rise. Turn into an assurance function and the time to approved production rises while exceptions pile up before anything ever ships. You can cheat either number on its own. You cannot cheat both at once. That pair is the honesty check a single scorecard could never be.
Legitimacy, not zero. The tempting metric here is wrong, and I know because I argued a version of it in public and was corrected. It is easy to say that any incident tracing back to a known, documented, accepted risk is a governance failure, and that the number should be zero. It is not. A bounded risk, accepted by someone with the authority to accept it, that then goes wrong, is not governance failing. It is governance working. Demand zero and the acceptances do not disappear, they go underground, which is worse. The real question is whether the acceptance was legitimate. Did the person have the authority. Was it documented before the event, not reconstructed after. Was the exposure bounded, and did the outcome stay inside that bound. Was it revisited once it did. A legitimate acceptance that goes wrong is a priced risk. An improper override is an unpriced one. Same incident, entirely different failure.
Look at when, not how many. The tell for find a way to yes is not the count of risk acceptances. It is their timing. If acceptances cluster in the days before a go live date, nobody is exercising judgment. They are clearing a deadline. Watch the distribution, not the total.
The simplest red flag. If the person has never said no, they are not doing the job. A governance lead who has blocked nothing is not evidence of a well designed pipeline. It is a rubber stamp with a job title. The first no is the proof the role is real.
The question behind the question
Notice what none of this measures. It does not measure how good the models are, or how fast they ship. It measures whether the authority to say stop is real, and whether the evidence behind a yes would survive being examined.
That is the whole job. Not to approve, and not to obstruct, but to make each yes one that an auditor, a regulator, and a board could all stand behind, and to have said no often enough that the yes still means something.
So the question is not whether you have an AI governance lead. That title is becoming standard furniture. The question is whether you could tell the difference between one who governs and one who signs. If you cannot, you do not have governance. You have a signature.