Compliance guide · Financial services
SR 11-7 & OCC 2011-12 for AI: a model risk management guide for banks.
The Federal Reserve’s SR 11-7 and the OCC’s 2011-12 bulletin have governed model risk management in U.S. banking for over a decade. Large language models, agents, and generative AI systems fall squarely within their definition of a “model.” This guide translates those traditional supervisory expectations into concrete AI governance controls for banking and financial services leaders.
Free download
SR 11-7 & OCC 2011-12 AI compliance checklist (PDF)
An 8-section, printable self-assessment covering inventory, validation, monitoring, governance, third-party model risk, and audit readiness. Use it to benchmark your program before your next exam or board review.
Why SR 11-7 and OCC 2011-12 apply to AI
SR 11-7 defines a model as any quantitative method, system, or approach that applies statistical, economic, financial, or mathematical theories to transform input data into quantitative estimates. OCC 2011-12 is substantively identical guidance issued jointly. Large language models, embedding models, RAG pipelines, and agentic systems that influence customer decisions, credit, disclosures, or operations meet this definition. Regulators have repeatedly signaled that generative AI is in scope, and examiners are already asking banks to show how their existing MRM programs extend to LLMs and third-party foundation models.
The six pillars, from traditional MRM to AI governance
Model inventory & tiering
SR 11-7 requires a comprehensive inventory of all models, including those from vendors. For AI and LLM deployments, this means cataloging every foundation model, fine-tune, RAG pipeline, agent, and prompt template in production, with a materiality tier that drives the depth of validation and monitoring required.
Conceptual soundness & development documentation
Regulators expect documented rationale for design choices, data lineage, and known limitations. Translated to AI: document the model card, training data provenance, evaluation methodology, prompt engineering choices, guardrails, and failure modes. Track this in a system of record, not scattered notebooks.
Independent validation
OCC 2011-12 and SR 11-7 mandate effective challenge by a party independent of model development. For LLMs, independent validation covers benchmark performance, adversarial and jailbreak testing, bias and fairness evaluation, hallucination rates on domain data, and stability across model versions and providers.
Ongoing monitoring
Models drift. LLMs also change under you when a provider updates weights or routing. Establish continuous monitoring for output quality, drift, latency, cost, refusal and hallucination rates, and downstream business KPIs, with thresholds that trigger review, rollback, or revalidation.
Governance, roles & controls
A governance framework defines model risk appetite, roles and responsibilities across the three lines of defense, approval workflows, and escalation paths. For AI, extend this to prompt review, human-in-the-loop policies, agent action scopes, and data-handling controls for sensitive inputs and outputs.
Policies, standards & training
Written policies and standards make the framework auditable. Cover acceptable use of generative AI, third-party model risk, vendor due diligence, incident response, and mandatory training for developers, validators, and business owners. Refresh at least annually as the AI landscape shifts.
Mapping the regulation to modern AI controls
| SR 11-7 / OCC 2011-12 requirement | Applied to enterprise AI & LLMs |
|---|---|
| Model definition (SR 11-7 §III) | Explicitly include LLMs, embeddings, RAG systems, agents, and prompt-based decision logic as models, even when built on third-party APIs. |
| Model development & documentation | Model cards, prompt/version registries, evaluation harness results, data sheets, and reproducibility artifacts stored in a central MRM (model risk management) system. |
| Effective challenge / independent validation | Independent red team: adversarial prompts, jailbreaks, prompt injection, data exfiltration, bias probes, hallucination and grounding tests on domain evaluation sets. |
| Ongoing monitoring | Production telemetry on quality, drift, refusal, hallucination, PII leakage, latency, cost per outcome, and business KPIs, with alerting and periodic revalidation. |
| Governance & policies | Board-approved AI risk appetite, RACI across 1LOD/2LOD/3LOD, approval gates by tier, human-in-the-loop rules, and vendor/third-party AI due diligence. |
| Vendor & third-party models | Provider risk reviews for foundation model vendors, SOC 2, data handling, indemnification, model change notice, evaluation transparency, and exit plans. |
A pragmatic rollout plan
You do not need to rebuild your model risk program from scratch. The fastest path to defensibility is to extend the MRM framework you already have to cover generative AI and third-party models, starting with the highest-materiality use cases.
Stand up the inventory
Discover every AI use case in production and in flight. Score materiality using existing MRM tiering. Freeze new high-risk deployments until minimum controls exist.
Extend MRM policy to AI
Amend the model risk policy to explicitly cover generative AI, agents, and third-party foundation models. Define documentation, validation, and monitoring standards by tier.
Independent validation & monitoring
Stand up an independent AI validation function or partner. Deploy monitoring for hallucinations, drift, and safety metrics on the top-tier models first, then work down.
Operationalize & audit
Integrate AI risk into board reporting, internal audit plans, and regulatory exam readiness. Run periodic tabletop exercises for AI incidents and vendor model changes.
Common pitfalls we see in bank AI programs
- •Treating LLM deployments as “IT projects” instead of models, bypassing MRM approval and validation entirely.
- •Relying on vendor claims for foundation-model performance without independent, domain-specific evaluation.
- •No monitoring for silent model changes when a provider updates weights or routing behind the same API.
- •Documentation stored in notebooks, tickets, or Slack instead of a system of record examiners can review.
- •Agentic systems deployed with broad action scopes and no human-in-the-loop policy for high-impact decisions.
How the Kennedy AI Scorecard helps
Our free scorecard benchmarks your AI operating model against the same dimensions examiners probe, strategy, governance, data, technology, and people. It is the fastest way to identify where your existing MRM framework already covers AI, and where you have gaps to close before your next exam or audit.
This guide is provided for informational purposes only and does not constitute legal, regulatory, or compliance advice. Consult your regulators, legal counsel, and internal risk functions for guidance specific to your institution.