OSFI Guideline E-23 and AI: What Financial Institutions Must Have in Place by May 1, 2027

OSFI E-23model risk managementfinancial servicesAI governanceCanada

OSFI Guideline E-23 makes every AI model inside a federally regulated financial institution a governed asset by May 1, 2027 — and the institutions that will pass their first review are the ones building the file now. Finalized in September 2025, E-23 (Model Risk Management) extends full lifecycle governance to AI and machine-learning models: an enterprise-wide model inventory, risk ratings, validation proportionate to risk, ongoing monitoring, and named accountability from design through decommission. One clarification first, because it derails real compliance programs: the AI guideline is E-23, not B-15 — B-15 is climate risk. The practical challenge is that the fashionable way to adopt AI, third-party cloud APIs, is the hardest to evidence under E-23: you cannot pin, inspect, or fully validate a model someone else changes on their own schedule. Self-hosted open-weight models invert that — every artifact E-23 asks for is producible from systems you control.

What does E-23 actually require?

E-23 organizes model risk management around the lifecycle, and its expectations map cleanly onto concrete artifacts:

E-23 expectation What the examiner will look for Cloud API reality Self-hosted reality
Model inventory Every AI model, use, owner, and risk rating enumerated "Model" is a vendor endpoint that changes silently Named model + pinned version per use case
Risk rating Materiality- and complexity-based ratings Complexity unknown; vendor discloses little Architecture, weights, and configs known
Validation Independent evaluation before use, proportionate to risk Point-in-time tests of a moving target Reproducible eval harness against a frozen model
Change management Re-validation when the model changes Vendor updates arrive unannounced Upgrades happen when you decide; evals gate them
Monitoring Ongoing performance tracking against baseline Baseline drifts with vendor releases Stable baseline; drift means your data changed
Explainability & documentation Understanding of model behavior and limits Black box behind an API Full access to weights, logs, and outputs
Accountability Senior ownership of model risk Split across vendor contract terms Entirely in-house

None of this prohibits cloud AI. It prices it: every vendor model becomes a validation subject you cannot hold still, and every silent vendor update is an unmanaged model change in your inventory.

Why is version pinning the heart of the E-23 problem?

Validation only means something against a fixed artifact. When a cloud provider updates its model — new checkpoint, new safety layer, new decoding defaults — the thing your institution validated no longer exists, and E-23's change-management expectation triggers whether or not you were told. Frontier API vendors deprecate versions on their schedule, not your review calendar; 2026 added a sharper lesson when the June 11 US export order cut foreign access to Anthropic's top models with 48 hours' notice. A model-risk file whose subject can be revoked, replaced, or altered by a third party is structurally unstable.

A self-hosted open-weight model is the opposite: a checksummed static file. The GLM-5.2 build you validated in month one is bit-for-bit the build serving members in month eighteen, until your own governance process — evaluation harness, sign-off, documented migration — decides otherwise. That is E-23's lifecycle diagram implemented in infrastructure.

What does a defensible E-23 AI stack look like?

The pattern we deploy for financial institutions — banks, insurers, and credit unions building to the federal bar:

  1. Pinned open-weight models. GLM-5.2 (744B MoE, 40B active, MIT license) as the interactive workhorse — member-service copilots, document drafting, knowledge retrieval — with DeepSeek for cost-efficient batch analytics. Checksums recorded in the model inventory.
  2. An evaluation harness built from your real work. Validation sets drawn from actual member interactions, lending documents, and policies — rerunnable on demand, producing the before/after evidence E-23's validation and change-management sections contemplate.
  3. Full-lineage logging. Prompts, outputs, model version, and adapter version logged in-house — usable for monitoring, explainability requests, and internal audit without a vendor data-sharing negotiation.
  4. In-perimeter data. Member data, lending files, and internal analytics never leave the institution — which simultaneously simplifies the privacy file (PIPEDA; for Quebec institutions, Law 25's cross-border rules, covered in our Law 25 AI compliance guide) and removes the CLOUD Act's reach over a US provider.
  5. Governed upgrades. New open-weight releases arrive monthly; each candidate runs the harness, and migration is a documented, reversible decision. The maintenance retainer is the change-management process.

The economics happen to point the same direction: above roughly 2M tokens per day, self-hosting runs 60–85% cheaper than metered APIs — our on-premise cost guide has the tiers — and overnight batch analytics run at zero marginal cost.

What should institutions do in the next two quarters?

May 1, 2027 is close enough that sequencing matters. The critical path we run with clients: inventory first (including the shadow-AI usage that surveys suggest roughly a quarter of employees already engage in), risk-rate the real use cases, then stand up the governed deployment — a six-to-twelve-week build via our on-premise LLM deployment practice — so that validation, monitoring, and documentation accumulate for months before the guideline bites. Institutions that treat E-23 as a paperwork deadline will spend 2027 retrofitting governance onto tools they cannot control. Institutions that treat it as an architecture decision will walk into their first review with an inventory of pinned models, a rerunnable harness, and logs that answer every question — which is what "model risk management" was supposed to mean all along.

Questions we get

Frequently asked questions

What is OSFI Guideline E-23?

E-23 is OSFI's Model Risk Management guideline — finalized September 2025, effective May 1, 2027 — governing how federally regulated financial institutions manage risk from models, explicitly including AI and machine-learning models. It requires an enterprise model inventory, risk-based ratings, validation proportionate to risk, ongoing monitoring, and clear accountability across the model lifecycle.

Does OSFI E-23 apply to generative AI and LLM copilots?

Yes. E-23's definition of a model is broad — methods applying theoretical, empirical, or statistical techniques to process data into outputs — and OSFI has been explicit that AI/ML models are in scope. A member-service copilot, a document-processing pipeline, or an internal LLM assistant used in decisions falls within the framework and must appear in the inventory with a risk rating.

Is B-15 the OSFI guideline for AI?

No — that is a common mix-up. Guideline B-15 addresses climate risk management. The AI and model-risk guideline is E-23 (Model Risk Management). AI governance programs citing B-15 are citing the wrong instrument.

Can a bank use ChatGPT or cloud AI APIs under E-23?

E-23 does not name vendors or prohibit tools; it requires lifecycle risk management the institution can evidence. That is the difficulty: a cloud API's model version, training changes, and behavior drift are outside the institution's control and often outside its visibility, so validation and change-management evidence is hard to produce. Self-hosted models with pinned versions make the same evidence straightforward.

Take the 40 Claude skills and the briefing with you

The Vault 2026 skills pack (calendar audits, hiring scorecards, calibration, continuity plans) plus the sovereignty briefing: model releases, deployment economics and regulatory shifts for regulated firms. One click to unsubscribe.

Free. You get the Vault 2026 skills pack now and the sovereignty briefing roughly monthly. One-click unsubscribe.

Ready to move from reading to running?

We design, build, fine-tune, host, and maintain sovereign AI deployments end to end.

Book a sovereignty assessment How deployment works