Bank Examination Evidence That Satisfies Model Risk Review

Key Takeaways

  1. Examiners accept contemporaneous, attributable, complete, and verifiable records that prove controls operated, not policies alone.
  2. For model and AI systems, that proof spans inventory, developmental evidence, independent validation with effective challenge, ongoing monitoring, and closed findings.
  3. Weak packs recycle undated screenshots, favorable samples without population logic, and validation that never challenges design assumptions.
  4. Agentic and generative systems still need the same evidence discipline, with decision lineage, policy verdicts, and change control filling surfaces traditional scorecards never had.

What Bank Examination Teams Reconstruct in Model Risk Review

A bank examination evaluates whether the institution operates safely, manages risk effectively, and complies with applicable law. When quantitative systems drive material credit, capital, liquidity, fraud, or customer decisions, model risk sits inside that review. OCC materials describe this work as testing management processes so banks avoid excessive risk and stay compliant [1].

Examiners Reconstruct Decisions, Not Slogans

In a bank examination, staff do not grade slogans. They reconstruct decisions. They ask who approved a model or AI application, when it entered production, and what follow-through occurred after change. The Federal Reserve Commercial Bank Examination Manual frames safety-and-soundness work around management activities and internal controls with a reconstructable trail [2]. That is the same reconstructability standard production AI observability is built to support.

Supervisory Guidance Evolution: From SR 11-7 to SR 26-2

For more than a decade, Supervisory Guidance on Model Risk Management organized expectations around development and use, validation, and governance. That body of guidance was known historically as SR 11-7 and OCC Bulletin 2011-12. In April 2026, the agencies issued revised interagency guidance through Federal Reserve SR 26-2 and OCC Bulletin 2026-13. The package clarifies risk-based model risk management, preserves the development-validation-governance architecture, and replaces the 2011 letters as the primary supervisory text [3].

What FDIC Examination Modules Sample

FDIC examination modules still show what staff sample when models are in scope:

  • Policies and inventory completeness
  • Developmental documentation
  • Validation reports with effective challenge
  • Ongoing monitoring and change control
  • Board or committee oversight
  • Vendor model practices

The winning answer to what evidence satisfies bank examiners is evidence quality plus lifecycle coverage. A thicker policy binder alone does not win a bank examination. Teams supporting AI in financial services should prepare for that reconstructability standard, even when a system sits outside a narrow model definition.

Evidence Quality Attributes Examiners Prefer Over Assertions

FDIC consumer compliance documentation standards illustrate the principle well: files should show a clear trail of decisions and supporting logic so a later reader can reconstruct the process [4]. Banks should hold themselves to a parallel bar, producing examiner-ready evidence that is requirement-linked, contemporaneous, attributable, complete, verifiable, and outcome-focused.

Requirement-linked artifacts map to a law, supervisory expectation, internal policy, or board standard, while contemporaneous records are created in the course of work, not rebuilt during exam week. Attributable records name preparer, reviewer, approver, and escalation path, and complete packs cover the review period and population, or document sampling methodology with enough rigor to reperform. Verifiable sources are system-backed so examiners can reperform against production truth, and outcome-focused evidence shows detection, ownership, remediation, and independent closure, not only design intent.

Weak evidence fails those tests in predictable ways: undated screenshots and self-attestations for material controls do not prove operation, favorable samples without population logic assert control without proving it, findings closed without retest leave gaps, and board minutes that only list topics lack substance. Policies without operating records rarely survive transaction testing.

For AI systems the same attributes apply to different artifacts: timestamped traces, evaluation scores, and guardrail allow, block, and redact verdicts replace the loan-file photocopy, while versioned prompts, configs, and inventory owners complete the set. Continuous Monitoring and full-fidelity telemetry produce the contemporaneous trail examiners reperform, because sampling-only evaluation weakens completeness — the unobserved path is where material exceptions hide. Build audit-ready logging before the notice letter arrives.

Assemble Bank Examination Evidence That Proves Controls Operated

Use the packet below as the exam-room checklist for material models and AI applications. Align each layer to your internal model risk management (MRM) or AI risk policy. Then prove the control ran for the period under review.

Establish Ownership and Inventory Coverage Before Examiners Ask

Start with a board-approved model risk or AI risk policy that defines scope, risk tiering, roles, escalation, and vendor standards. Pair it with a complete inventory of systems in use, under development, and retired.

Each entry should carry owner, purpose, risk tier, last and next validation or review dates, known limitations, and material dependencies. Independence of validators must be real in reporting lines, not paper-only separation.

Enterprise Auditable Governance that includes a live AI registry of models and applications makes this layer maintainable instead of a pre-exam scramble.

Document Design Decisions Early to Defend Intended Use

Developmental evidence states purpose, design logic, data provenance, assumptions, alternatives considered, and known limitations before production use. For machine learning, generative systems, and agents, document feature or prompt design and pipeline boundaries.

Include retrieval, tools, system prompts, version pins, and environment parity checks. Intended-use constraints and prohibited uses belong in the file before go-live.

Do not rebuild them as a retrospective memo after an incident. The revised supervisory attachment continues to stress documentation and developmental evidence quality as foundations for validation and monitoring [5].

Use Independent Challenge to Keep Weak Validation Out of Production

Validation should cover conceptual soundness, implementation verification, and outcomes analysis. Scale benchmarking or sensitivity work to risk tier.

Written challenge matters. Raise findings, assign severity, state conditions for use, and honor fail grades rather than soften them to keep a model live.

Vendor and foundation models require institutional outcomes testing, customization assessment, and conceptual understanding of limitations. Pass-through vendor materials without bank-specific testing are not validation.

Maintain a Monitoring Trail That Makes Remediation Verifiable

Ongoing monitoring must show performance or behavioral checks with thresholds and escalation. An annual memo is not enough for high-change systems.

Change tickets should record what changed, who approved, and which revalidation triggers fired. Issue logs need owner, due date, fix evidence, and independent confirmation of closure.

Board or committee management information should show risk profile, material exceptions, and overdue findings rather than topic lists. Reliable Evaluation that uses the same standards in pre-production and production shortens assembly of this layer because the monitoring trail already exists.

How Agentic Systems Change Bank Examination Evidence Surfaces

Why Traditional Model Assumptions Break Down

Traditional model assumptions include relatively stable inputs, a fixed estimator, and reproducible outputs. Generative and agentic systems break those assumptions. The unit of review is often a system: prompt, retrieval, tools, model, and downstream action. Non-determinism, tool use, and third-party components expand the failure surface that governance must document.

Regulatory Scope and Examiner Expectations

The April 2026 revised model risk guidance states that generative AI and agentic AI models are novel and rapidly evolving. It places them outside the formal scope of that guidance [6]. That scope note does not erase examiner interest. Safety-and-soundness, consumer compliance, third-party risk, and operational resilience still demand reconstructable proof when agents act on capital, customer data, or developer environments.

Compensating Evidence for Agentic Systems

Compensating evidence fills the surfaces traditional scorecards never had. Key practices include:

  • Behavioral and adversarial testing plus human-in-the-loop controls for high-impact outputs
  • Decision lineage captured across tool calls and runtime policy enforcement records that show governance operated on the request path
  • Expanded inventory covering first-party agents your teams build, third-party agents you deploy, and coding agents your developers use: all need owners and controls where they touch enterprise data

Agentic Observability that joins agent-side telemetry with gateway capture creates examiner-visible lineage instead of post-incident reconstruction. That is Standardized Telemetry applied to the agentic stack.

Concrete Example: Loan-Servicing Agent Evidence Trail

Consider a concrete loan-servicing agent request. The agent pulls a customer note through a tool call, drafts a balance explanation, and attempts to post a fee waiver. Examiner-ready proof for that single path includes:

  • The session trace
  • Tool inputs and outputs
  • The Enforceable Policy verdict that redacted a primary account number
  • The human approval ticket for the waiver
  • The closed monitoring alert if the agent later retried the same action out of policy

Inline policy enforcement records matter here. Allow, block, and redact verdicts with timestamps, policy version, and actor context prove that AI Guardrails ran when the request happened.

In-Environment Evaluation Without External Dependencies

When evaluation must stay inside the bank's environment, Fiddler Centor Models (formerly Fiddler Trust Models) lead with a batteries-included, in-environment design. No external LLM call is required to evaluate an agent or LLM output. Evaluations run inside the customer's own environment. No data leaves, no external API is called, and no per-evaluation cost is incurred. They also deliver under 80ms response time and remain framework, model, and cloud agnostic across stacks such as Azure OpenAI, Amazon Bedrock, LangGraph, and Google Gemini.

Examiners still want the same quality attributes. The artifacts change from scorecards to traces, verdicts, and session context.

Common Evidence Failures That Trigger Follow-Up Findings

Before the bank examination window opens, self-audit against the failures examiners escalate most often:

  1. Incomplete inventory: Shadow models or agents in production without owners, tiers, or next-review dates.
  2. Thin developmental files: Purpose stated without data provenance, assumptions, limitations, or alternatives considered.
  3. Validation without challenge: Reports that affirm design without independent findings, conditions, or fail authority.
  4. Annual-only monitoring: High-change AI systems reviewed once a year while prompts, tools, and models churn weekly.
  5. Vendor pass-through: Foundation or vendor models accepted on marketing materials without institution outcomes testing.
  6. Closed without retest: Issues marked complete without independent confirmation against production evidence.
  7. Board packs without substance: Approvals recorded without risk discussion, exception trends, or overdue finding aging.

Any one of these can turn a clean narrative into a matter requiring attention. Fix inventory and monitoring first, since they're the fastest way to restore completeness and contemporaneity across the packet.

They are the fastest way to restore completeness and contemporaneity across the packet. Patterns such as anomaly detection in financial agent workflows only help if alerts, ownership, and closure land in the same auditable trail.

Conclusion: Build the Trail Before the Bank Examination Window Opens

Successful bank examination outcomes rest on operable, reconstructable proof across the model and AI lifecycle. Policies set expectations.

FDIC consumer compliance examination documentation standards put it, evidence should let a later reviewer reconstruct decisions and supporting logic from the file itself [4]. That means showing the control ran, who owned it, what broke, and how the bank fixed and rechecked it against source systems.

Practical next step for the next bank examination cycle: pick one material AI or model system this week and assemble the full evidence packet above. If inventory or ongoing monitoring is incomplete, remediate those layers first so every other artifact has a population and a time base.

For teams that need Reliable Evaluation, Continuous Monitoring, and Auditable Governance trails for agents and models, request a demo of the Fiddler AI Observability and Security Platform.

References

[1] Office of the Comptroller of the Currency, "Examinations Overview," OCC. [Online]. Available: https://www.occ.gov/topics/supervision-and-examination/examinations/examinations-overview/index-examinations-overview.html

[2] Board of Governors of the Federal Reserve System, "Commercial Bank Examination Manual," Federal Reserve, updated Feb. 19, 2026. [Online]. Available: https://www.federalreserve.gov/publications/supervision_cbem.htm

[3] Board of Governors of the Federal Reserve System, "SR 26-2: Revised Guidance on Model Risk Management," Federal Reserve, Apr. 17, 2026. [Online]. Available: https://www.federalreserve.gov/supervisionreg/srletters/SR2602.htm

[4] Federal Deposit Insurance Corporation, "II-7 Documenting the Examination," FDIC Consumer Compliance Examination Manual, updated Aug. 2026. [Online]. Available: https://www.fdic.gov/consumer-compliance-examination-manual/ii-7-documenting-examination

[5] Board of Governors of the Federal Reserve System et al., "Supervisory Guidance on Model Risk Management," SR 26-2 attachment, Apr. 17, 2026. [Online]. Available: https://www.federalreserve.gov/supervisionreg/srletters/SR2602a1.pdf

[6] Office of the Comptroller of the Currency, "OCC Bulletin 2026-13: Model Risk Management: Revised Guidance," OCC, Apr. 17, 2026. [Online]. Available: https://www.occ.gov/news-issuances/bulletins/2026/bulletin-2026-13.html

Frequently Asked Questions

What Evidence Satisfies Bank Examiners for Model Risk?

Contemporaneous inventory, developmental documentation, independent validation with effective challenge, ongoing monitoring logs, change approvals, and closed findings mapped to policy. Examiners reperform against source systems. Screenshots and self-attestations without attributable operators rarely suffice. Build the trail during ordinary operations, not during bank examination week.

Does SR 11-7 Still Apply to Machine Learning and Generative AI?

SR 26-2 and OCC Bulletin 2026-13 supersede SR 11-7 as the primary interagency model risk text. They keep a risk-based development, validation, and governance architecture for models in scope. Generative and agentic AI sit outside that revised guidance's formal scope. Institutions still adapt testing and documentation for opacity, drift, non-determinism, and third-party components under broader safety-and-soundness expectations.

What Is the Difference Between a Policy and Examiner-Ready Evidence?

A policy states expected behavior, roles, and standards. Examiner-ready evidence shows the control ran for the period under review. It shows who owned each step, what exceptions occurred, and how issues were fixed and independently rechecked. Without operating records tied to systems of record, policy text is assertion, not proof.

How Often Should AI Models Be Monitored Between Validations?

Frequency should match risk tier and change rate. Material or rapidly changing AI systems need recurring performance and behavioral checks with thresholds, escalation paths, and retained logs. Annual validation memos alone are a poor fit when prompts, tools, retrieval corpora, or model versions change continuously. Pair scheduled Reliable Evaluation with production monitoring so the evidence period matches real change velocity.