Your AI Literacy Program Measures Attendance
Most AI literacy programs in regulated institutions teach what the technology is and then count who turned up. That number goes into a board pack and satisfies a supervisor asking whether staff have been made aware. It tells you nothing about the only literacy that governs anything: whether the person who can approve an AI system can evaluate a claim made about it.
Those are different competences and only one is load bearing. You cannot catch the gap with a quiz on what a transformer is, because the people who fail the real test would pass that one comfortably.
The law has moved on this, and in the wrong direction. Article 4 of the EU AI Act has applied since February 2025, and as first written it asked providers and deployers to ensure, to their best extent, a sufficient level of AI literacy among their staff. Since the AI Omnibus entered into force in July 2026, it asks them instead to take measures that support the development of AI literacy. The obligation moved from an outcome to an activity, which is to say toward attendance. The legal floor now asks whether you took measures, not whether anyone who signs can evaluate a claim. That makes the internal measure more important, not less: the only test of the competence that governs will be the one you run yourself.
A true principle, applied to the wrong layer
Here is an objection I have been given more than once, by senior people, in rooms where an agent proposal was on the table: generative AI is non-deterministic, and compliance requires deterministic outputs.
It deserves to be taken seriously. It uses the right words, it is made in good faith by people doing the job the institution pays them for, and it rests on a principle that is true, since a control you cannot reproduce is not a control. It is also wrong, in a specific way, and the specificity is the whole point.
Determinism is not a property you inherit from a component. It is a property you engineer into a system. Fix the procedure before any model touches it, ground the retrieval, let the model interpret and never decide, log every step, and have a human review the record rather than repeat the work. That is how I have built it and put it into production, inside the compliance function of a global systemically important bank, with every function that could have blocked it in the room from the first week. I met the objection most often on other proposals, in other rooms across financial services.
Notice what became deterministic and what did not. The procedure is fixed, the evidence is complete, and the decision is a human's, recorded. The model's wording is not reproducible and never needs to be, because nothing in the control depends on it. Determinism was not removed from the system. It was moved to the layer where the control actually sits.
The same objection is right about shadow AI. A model somebody runs privately has no fixed procedure, no log and no recorded decision, so determinism was never engineered at any layer. Telling those two cases apart is the competence this piece is about.
The gap in those rooms was not ignorance. The objection was fluent, correctly worded, and delivered with justified confidence by people with every reason to be confident about compliance. A true principle was applied to the wrong layer, and not one person in the room could tell, including, before I had built the thing, me. That is invisible from every measure an institution currently uses, which is why it keeps happening.
A no that nobody can reconstruct teaches nothing
The second failure leaves less evidence.
I have watched it from the proposing side, and the honest account is that I could not tell you afterward what the decision turned on. Neither, I suspect, could the people who made it. That is not a complaint about the outcome, which may well have been right. It is the problem itself. An institution that cannot name the standard it applied cannot tell a disciplined no from an accidental one, learns nothing from either, and repeats whichever it happened to produce.
Plenty of AI proposals deserve to die, and a firm that declines one for a reason it can state and defend is working properly. The failure is not the no. It is being unable to tell which kind you just said.
Why this sits with governance and not with training
Here is the claim, and it is the one I expect disagreement on.
AI literacy is a governance control, not a development activity, and filing it under learning is a category error the org chart makes for you. Training functions are measured on reach and completion, correctly, because that is what training functions are for. Put literacy there and it gets optimized toward those measures, which are exactly the measures that cannot detect the failure.
The strongest objection is that governance has no curriculum. Risk and compliance do not run capability building, they have no pedagogy and no training budget, and moving literacy to them puts a control in a function that cannot deliver it. That objection is right about the teaching and wrong about the control. Let the training function keep the classroom. What has to move is the measure and the ownership of it: who decides what counts as sufficient, and who answers when it is not.
This is the scorecard problem again, with the same remedy. I have argued that no single function should own the scorecard for a governance role, because a metric drifts toward whoever writes it. A literacy bar is no exception. Owned by training, it drifts toward completion. Owned by governance alone, it drifts toward caution, because caution costs the people who set it nothing. Owned by delivery, it drifts toward whatever lets this week's decision through. So governance sets what sufficient means, and the function that has to live with the pace countersigns it.
What replaces the attendance figure
Four parts, and none of them is a course.
Name the decisions. Which approvals this person actually holds. Your delegation of authority already says.
Name what would have to be true. For each of those, the two or three claims that carry it, the ones that would make the approval wrong if they failed.
Demonstrate rather than attest. Put a real proposal in front of them carrying a layer error, the kind determinism was, and see whether they find it. Someone who finds it has shown the competence. Someone who does not has shown you where the exposure sits, which is the more useful result.
Record it where authority lives, next to the delegation and not in the learning system, because nobody consults the learning system on the day something gets signed.
Then two numbers go to the board, and the second must not be one you author. The share of named approvers who have demonstrated: governance produces that. Against it, the share of approvals later found to have rested on a claim that failed: reality produces that one, after the fact, out of decisions already taken. Soften the exercise and the first number rises while the second does not move, because the exercise does not generate it. A pair only works when one half is out of your hands.
Two objections arrive together here, and they are the same objection. Senior people will not sit an examination, and once the planted errors are known the thing becomes theater. Both are right about a test run as a test. Neither touches what I am describing, which is a decision. You are not examining the approver. You are putting a real proposal in front of them, which is what they were going to do anyway. The only difference is that somebody already knows where the weak joint is.
None of which means your people need to become engineers, or that the awareness training was wasted. Awareness is the floor: someone who does not know what a model is cannot evaluate anything, so that spending was necessary and remains so. The problem is only that it has been measured as though it were the whole building.
The test you can run this quarter
Do not run a survey.
Take the last three AI proposals your institution declined, or quietly let die. For each one, write down the reason it was actually judged against, in one sentence, in the words used at the time.
Then ask the only question that matters: could anyone in that room have tested that reason?
Three answers come back and each is a different finding. The reason was tested and sound, so take the confidence. Or it was a true principle applied to the wrong layer, the way determinism was, and sounded rigorous enough that nobody challenged it. Or nobody can reconstruct a reason at all.
Then run it on the proposals you approved. That half is less comfortable and tells you more: an institution that cannot test a bad objection cannot test a good pitch either.
If you cannot do this because nothing was written down, stop there. That is the finding, and it is the same answer as the third case.
One caution, because your general counsel will reach it before your second proposal. Writing down why AI decisions were actually taken creates a record of your own governance, in your own hand, discoverable in examination or litigation. Run it where such records already belong, in committee minutes or under privilege.
The position I hold
AI literacy at the decision level is not knowledge of the technology. It is the ability to evaluate a claim about it, and to recognize the edge of your own judgment. Governance asks who is accountable when a system decides. Literacy asks whether that person can tell whether they should have said yes. Governance enables trust, and trust enables speed, but neither can be extended to a judgment nobody in the room is able to make.
Measure attendance and you learn who turned up. An institution that has trained its staff and filed the certificates holds a delivery record, not a custody record. It can show the training was handed over. It cannot show that anyone is holding the capability now.