RemKey

Guides

What do OSFI B-13 and E-23 mean for a firm's LLM usage?

OSFI's technology-risk guideline (B-13) and revised model risk guideline (E-23) both reach AI use at federally regulated financial institutions. What each asks for at a high level, and the evidence a firm needs to have ready.

First, the disclaimer that matters: this is a plain-language orientation, not legal or regulatory advice, and we are neither your counsel nor OSFI. Read the guidelines themselves and talk to your compliance function.

Two OSFI guidelines frame how federally regulated financial institutions in Canada are expected to manage the risks that LLM usage creates. Neither says "LLM" on every page. Both reach LLM usage anyway.

B-13: technology and cyber risk

Guideline B-13 (in effect since January 2024) sets expectations for technology and cyber risk management: knowing your technology assets, managing third-party technology dependencies, and being able to demonstrate sound risk management of the technology your business runs on. LLM API usage is squarely that: an external technology dependency, carrying data out of your environment, embedded in business workflows. The uncomfortable B-13 question for most firms is not "is your AI vendor secure," it is "do you know everywhere your teams call an LLM at all?" Ad-hoc scripts and developer tools pointed straight at a provider are technology assets nobody inventoried. Finding them is the first job; our note on shadow AI covers how.

E-23: model risk, explicitly including AI

OSFI's revised E-23 guideline on model risk management explicitly broadens "model" to capture AI and machine-learning models, with lifecycle expectations: inventory, documentation of what a model is used for, monitoring, and accountability. The revised guideline takes effect May 1, 2027, which sounds far away and is not: an inventory-and-evidence discipline is much cheaper to start while usage is small than to reconstruct retroactively across two years of ungoverned calls. If your firm uses foundation models through APIs, the practical E-23-shaped questions are: which models, for which use cases, with what screening, and where is the record that shows it.

The common denominator: evidence, per call

Strip the two guidelines to what an examiner can actually sample and both converge on the same artifact: a complete, trustworthy record of what your AI usage did. Which model handled a request and why. Whether customer data was screened before leaving. Whether the record itself can be trusted, or merely believed. That artifact is what RemKey produces automatically: every call through one base URL is screened fail-closed, routed with a documented decision, and written to a hash-chained, Ed25519-signed ledger your auditor verifies offline. The financial-services page maps each mechanism to the questions compliance teams get; the trust page covers how we govern ourselves, including what we do not guarantee.

Want the answer applied to your stack? Start free and the evidence trail begins with your first request, or read the other guides.