At some point in every AI governance program, an executive asks the question that the entire program has been implicitly avoiding: “So — are we in good shape or not?”

The honest answer is usually complicated. The answer that gets given is usually a framework diagram. And the answer that would actually be useful is four or five numbers that a non-specialist can interpret in ninety seconds.

Building those numbers is harder than it looks, because most of what’s easy to count in AI governance is activity, and activity is not risk. Having briefed CIOs, VPs and risk committees for a couple of decades — and having watched plenty of good programs lose funding because they reported effort instead of exposure — here’s how I’d construct the measurement layer.

The rule that governs everything else

Report exposure and coverage, not effort.

“We reviewed 34 AI use cases this quarter” is effort. It sounds productive and it answers nothing — 34 out of what? An executive cannot tell from that number whether the situation is improving or deteriorating.

“91% of known AI systems have an assigned owner and risk tier, up from 64% last quarter; 3 high-risk systems are running without a current impact assessment” is exposure. It has a denominator, a direction, and an implied action.

Every metric below has a denominator. That’s not stylistic — it’s the difference between a dashboard and a status report.

Four families of metric

1. Coverage — do we know what we have?

This is the foundational family, and in the first year it’s the only one that matters much, because every other metric is unreliable until coverage is high.

  • AI systems in the inventory, trended, split by risk tier.
  • Inventory completeness — the honest one. Estimate it by sampling: run a discovery sweep (CASB, expense data, SaaS management platform, procurement records) and measure what fraction of what you find was already in the inventory. If discovery finds 40 systems and 25 were known, coverage is roughly 63%, and you should say so.
  • Percentage of systems with a named business owner. Ownerless AI systems are the single most actionable finding you can put in front of an executive, because the fix is a phone call.
  • Percentage with a current risk tier.
  • Shadow AI signal — unsanctioned generative AI usage detected, trended. Sourced from your CASB and Defender for Cloud Apps telemetry.

2. Timeliness — is governance a bottleneck?

The family that governance functions least want to measure and most need to. If your process is slow, the business routes around it, and every other metric degrades as a consequence.

  • Median time from intake to decision, split by risk tier. Track the median and the 90th percentile; the tail is where the reputational damage happens.
  • Percentage of decisions within SLA.
  • Open queue age — how long the oldest pending item has been waiting.
  • Assessments overdue for periodic review.

I’d put at least one timeliness metric on the executive page, deliberately. It signals that the governance function holds itself to a service standard, and it’s the metric most likely to earn you resourcing.

3. Risk posture — what’s our actual exposure?

  • High-tier systems in production without a completed impact assessment. Target: zero. Any non-zero number is the headline.
  • Systems in production without approved monitoring.
  • Open accepted residual risks, with count past expiry. Accepted risk with no expiry date isn’t accepted, it’s forgotten — so measure the expiry breach explicitly.
  • AI-related incidents and near-misses, by category. Use a stable taxonomy; the NIST Generative AI Profile categories work well.
  • Third-party AI systems without completed due diligence.
  • Vendor model changes notified in the period, and how many triggered re-evaluation. This one is quietly excellent, because it makes visible a risk that changes without any action on your part.
  • Human override rate on decision-support systems, trended. A rising override rate is your earliest signal of model degradation, and it’s a metric a business executive intuitively understands.

4. Assurance — does anyone independent agree with us?

  • Percentage of high-tier systems with independent validation or audit coverage.
  • Open audit findings related to AI, by age.
  • Control testing results — tested, effective, exceptions.
  • Certification or conformance status if you’re pursuing ISO/IEC 42001 — nonconformities open and closed.

Family 4 is what converts your self-reported numbers into something a board can rely on. Without it, every other metric is management marking its own homework, and sophisticated boards know it.

The one-page executive dashboard

Structure I’d use. One page, and I mean one.

Top strip — four numbers, large. AI systems in inventory (with quarter-over-quarter delta). Percentage governed to standard for their tier. High-risk items requiring attention. AI incidents this period.

Left — the risk posture panel. A tier distribution bar (how many systems at each tier), and a short list of the specific items requiring executive attention. Name them. “Three high-tier systems in production without current impact assessments: [system], [system], [system] — owners engaged, target close date X.” Vagueness here reads as evasion.

Right — trend. Two or three lines over four to six quarters: coverage, systems governed to standard, incidents. Trend is what tells an executive whether the program is working. A single-period snapshot never can.

Bottom — one short narrative block. Three or four sentences: what changed, what’s driving it, what decision or support you need. This is the part that actually gets read. Write it last and write it in plain language.

What to leave off: framework diagrams, maturity radar charts with eleven axes, control counts, and anything requiring the reader to know what “MEASURE 2.11” means. Keep the detailed pack behind the summary for the people who want it — and they will be a small minority.

If you’re building the visual layer, resist the urge to make it dense. The most effective governance dashboard I’ve seen fit four numbers and one trend line on a slide, with the detail available on request. The one that got ignored had nineteen gauges.

Three metrics worth more than their weight

If you can only build three, build these:

Coverage of the inventory. Everything else is conditional on it, and the first honest measurement is usually the single most persuasive slide the program will ever produce.

High-tier systems out of compliance with their own tier requirements. It’s a small number, it’s actionable, and it maps directly to the question executives are actually asking.

Median time to decision. It’s the metric that keeps the governance function honest and keeps the business inside the process.

Building the narrative around the numbers

Metrics don’t persuade on their own. The structure that works with senior audiences:

  1. Position — where we are, in one sentence with a number.
  2. Direction — better or worse than last period, and why.
  3. Exposure — the specific things that would matter if they went wrong, named.
  4. Ask — the decision, funding, or executive sponsorship required.

Then stop. The most common failure in executive reporting isn’t insufficient detail — it’s a twelve-slide pack that buries the ask on slide eleven, by which point the meeting has moved on.

Two other habits worth adopting. Report the uncomfortable number. A dashboard that only ever shows green stops being read, and the first time something goes wrong, every previous green becomes retrospectively suspect. And keep the metric definitions stable. Changing how you calculate coverage between quarters destroys the trend line, which is the most valuable thing on the page. If a definition must change, restate the prior periods.

Where the numbers come from

A practical note, because this is where dashboards die. Every metric above should be a by-product of the workflow, not a data-collection exercise. If producing the quarterly dashboard takes a person a week of chasing spreadsheets, it will be produced twice and then quietly abandoned.

That means the inventory and lifecycle workflow has to be the system of record: intake writes the record, tiering updates it, the impact assessment attaches to it, the launch gate stamps it, monitoring feeds it, review refreshes it. Get that right and the dashboard is a query. Get it wrong and the dashboard is a project, every single quarter.

If you’re building AI governance reporting for an executive or board audience and want a second opinion on the metric set, get in touch.