Financial services organizations have an advantage in AI governance that they consistently fail to use: they’ve been governing consequential algorithms for decades.

Model risk management is a mature discipline in banking. There are model inventories, tiering schemes, independent validation functions, documented assumptions and limitations, ongoing performance monitoring, and a three-lines structure with real teeth — all of it built long before anyone was worried about a large language model. When a bank stands up a brand-new “AI governance framework” in parallel to that, it usually produces duplication, jurisdictional friction between the AI governance team and the model risk team, and a set of controls that are weaker than the ones already in the building.

Having spent much of my career in and around banking — IS audit at ADCB, information security risk across Barclays’ emerging markets operations, and banking and pension audit work at PwC covering SOX, CDIC and Volcker Rule engagements — my strong view is that the right move in a regulated financial institution is to extend model risk management to cover AI, and then identify precisely where it doesn’t reach. Both halves matter.

What model risk management already gives you

Map the AI governance requirements against an existing MRM framework and the overlap is substantial:

AI governance need MRM already provides
AI system inventory Model inventory, with ownership and tiering
Risk tiering Model tiering by materiality and consequence
Impact/risk assessment Model risk assessment and documented limitations
Pre-deployment evidence Independent model validation
Ongoing monitoring Model performance monitoring and back-testing
Change control Model change management and re-validation triggers
Independent challenge Second-line validation function
Assurance Internal audit coverage of the MRM framework
Documentation standards Model documentation requirements

Supervisory expectations reinforce this. In Canada, OSFI’s Guideline E-23 on model risk management is explicitly scoped beyond traditional quantitative financial models to the broader population of models including machine learning and AI, and it expects institution-wide governance rather than a framework confined to capital models. In the US, SR 11-7 / OCC 2011-12 remains the reference point for model risk management, and supervisors have been clear that AI-based models are models. The regulatory direction of travel is consistent: don’t invent a separate regime, apply the one you have and make sure it stretches.

That’s the message to take to your model risk committee. Not “AI needs a new framework” but “your framework’s scope definition needs to catch these, and here are the four places it currently can’t.”

Where model risk management genuinely doesn’t reach

This is the part that requires honesty, because the extension argument fails if you pretend the fit is perfect. Four gaps matter.

1. Non-decision AI. MRM exists to govern models that produce outputs used in business decisions. A large share of enterprise generative AI does something else: it drafts, summarizes, searches, and assists. A Copilot deployment across 40,000 employees makes no decisions and falls outside a strict model definition — while creating substantial data exposure, confidentiality and information-integrity risk. If you extend MRM naively, this entire category falls through the floor.

The fix is a scoping decision made deliberately: MRM governs AI in decisioning; a parallel track governs enterprise AI tooling, with the data-governance and information-security controls appropriate to it. Both feed one inventory and one reporting line.

2. Validation methodology. Traditional validation tests conceptual soundness, ongoing monitoring, and outcomes analysis against known-good benchmarks. Much of that doesn’t transfer cleanly to a foundation model you didn’t train, can’t inspect, and whose provider may change under you. Validating a generative system means evaluation against task-specific criteria, adversarial and red-team testing, and monitoring for behaviours — confabulation, prompt injection susceptibility, harmful output — that have no analogue in a credit scorecard. Your validation function needs new methods and, frankly, new skills. Assuming existing validators can absorb this without investment is the most common execution failure I’d expect.

3. Third-party model dependency. MRM frameworks handle vendor models, but usually vendor models that are static, documented and licensed. A foundation model accessed through an API can change materially without notice, without a version bump you can detect, and without any change to your own environment. That breaks the assumption underneath re-validation triggers. You need contractual model-change notification, independent periodic re-evaluation regardless of notification, and a vendor risk process that treats the model provider as a critical dependency rather than a software supplier.

4. Agentic AI. MRM has no concept of a model that takes actions — calls tools, writes to systems, initiates transactions, chains steps autonomously. That’s an operational and access-control risk problem wearing a model’s clothing, and it needs the agentic governance controls that MRM was never designed to provide: identity and least privilege for non-human actors, connector and tool governance, action logging, and hard limits on autonomy in regulated processes.

Where the frameworks fit

Financial institutions carry a lot of framework overhead already, so be precise about what each thing is for:

  • MRM (E-23, SR 11-7) — the supervisory expectation. Non-negotiable, and the natural home for AI in decisioning.
  • NIST AI RMF — the risk vocabulary and structure, particularly the trustworthiness characteristics and the generative AI risk taxonomy. Use it to enrich your MRM risk assessment, not to replace it.
  • ISO/IEC 42001 — the certifiable management system. Most valuable to financial institutions that sell AI-enabled services and need to demonstrate governance to counterparties, or that want an external certificate alongside internal supervisory compliance.
  • Existing operational risk, conduct risk, privacy and third-party frameworks — where a lot of AI risk actually belongs, and where it will get sustained attention because those frameworks are already funded.

The unifying principle: AI risk should not have its own risk taxonomy. Map it into the enterprise taxonomy — operational, conduct, compliance, model, third-party, information security, privacy — so it competes for attention on the same terms as everything else. Risk categories that exist only because a technology is fashionable get defunded when the fashion changes.

The conduct and fairness dimension

One thing financial services must handle that many sectors can defer: consequential decisions about people, at scale, in a regulated context. Credit adjudication, underwriting, pricing, fraud and AML alerting, collections prioritization, and increasingly hiring. These carry fair-lending, human-rights-code, consumer-protection and conduct-risk exposure that a general AI governance framework treats as one bullet among many.

Practically, high-tier decisioning systems affecting individuals need: fairness testing against defined protected characteristics with documented methodology; explainability sufficient to give an affected person a reason, not just a score; a documented human review path with genuine authority to override; monitoring for disparate outcomes over time, not only at validation; and retained records long enough to answer a complaint or a regulatory inquiry years later.

That last point interacts awkwardly with model retirement. When a model is decommissioned, the records of the decisions it made about people generally cannot be. Build that into your decommissioning process rather than discovering it during a complaint investigation.

Three lines, without the turf war

The structural risk in a financial institution isn’t a missing framework — it’s overlapping ones. AI governance, model risk, operational risk, technology risk, privacy and information security all have a legitimate claim on AI, and each will build its own intake if allowed to.

What resolves it:

  • One intake, one inventory. Multiple assessment tracks are fine; multiple front doors are not. The business should answer the AI question once.
  • A published scoping rule that says which track a system follows, decided at triage. Ambiguity here becomes a standing agenda item that never resolves.
  • Clear second-line ownership of the aggregate view, so someone can answer “what’s our total AI exposure” without a three-week reconciliation.
  • Third line independent of the design. Internal audit should be reviewing the AI governance framework for auditability early — but not co-authoring the controls it will later test.

None of this is novel governance thinking. That’s the point. Financial services organizations have spent thirty years learning how to govern consequential algorithms under supervisory scrutiny, and the institutions that do best with AI are the ones that recognize what they already have — then invest specifically in the four places it genuinely falls short.

Supervisory guidance in this area continues to evolve; confirm current expectations and effective dates with your regulator or compliance function.

If you’re extending model risk management to cover AI, or working out where the boundary between MRM and enterprise AI governance should sit, let’s talk.