Executive summary
- The current version of PRA supervisory statement SS1/23 on model risk management principles for banks was published and effective on 23 April 2026.
- SS1/23 is structured around five principles: model identification and model risk classification; governance; model development, implementation and use; independent model validation; and model risk mitigants.
- The PRA’s 2025/26 annual report states that it completed the initial review of the first cohort of firms and plans to conclude reviews of the remaining firms by the end of 2026, including discussion with chief risk officers on AI and machine-learning models.
- The EBA’s June 2026 Risk Assessment Report notes that frontier AI can amplify cyber risk and that banks need risk-sensitive, transparent AI governance supported by data and cyber security and by testing.
- The EBA also confirms that requirements under the Digital Operational Resilience Act apply to banks’ use of AI systems, so AI model risk and ICT risk management cannot be run as separate programmes.
- Practical 2026 priorities are a complete AI/ML inventory with defensible tiering, explainability and limitation statements, drift monitoring, vendor model transparency, evidenced human oversight and board-level accountability.
The 2026 regulatory position
Model risk management for UK banks is governed by the Prudential Regulation Authority’s supervisory statement SS1/23, Model risk management principles for banks. The version currently in force was published and took effect on 23 April 2026. It applies model risk expectations across a firm’s model estate rather than only to regulatory capital models, which is the reason artificial-intelligence and machine-learning systems fall within scope wherever they meet the statement’s definition of a model.
Supervisory attention has moved from framework design to implementation evidence. The PRA’s 2025/26 annual report states that it completed the initial review of the first cohort of firms and plans to conclude reviews of the remaining firms by the end of 2026, and that these reviews include discussion with chief risk officers on artificial-intelligence and machine-learning models. Firms should therefore expect their AI inventory, tiering rationale and validation evidence to be examined directly rather than described.
At European level, the European Banking Authority’s June 2026 Risk Assessment Report observes that frontier artificial intelligence can amplify cyber risk, that banks need risk-sensitive and transparent governance of AI use, and that data security, cyber security and testing are central to safe adoption. The report also confirms that requirements under the Digital Operational Resilience Act apply to banks’ use of AI systems.
The five SS1/23 principles applied to AI
SS1/23 is organised around five principles. Each translates into concrete obligations once machine-learning systems enter the estate.
- Model identification and model risk classification: the firm maintains a complete inventory of models, with a definition wide enough to capture quantitative methods embedded in tools, spreadsheets and vendor products, and classifies each by model risk.
- Governance: the board and senior management own the framework, roles are separated, and accountability for model risk is assigned to identified individuals rather than to committees alone.
- Model development, implementation and use: development is documented, testing is proportionate to risk, and use is confined to the purpose for which the model was approved.
- Independent model validation: a function independent of development performs risk-based validation, challenges conceptual soundness and outcomes, and issues findings that are tracked to resolution.
- Model risk mitigants: where residual risk remains, post-model adjustments, use restrictions, overlays and enhanced monitoring are applied under explicit governance rather than informally.
None of these principles is AI-specific. What changes with AI is the evidence burden. A gradient-boosted credit-decision model and a logistic scorecard sit under the same principle on independent validation, but the machine-learning system requires substantially more work to demonstrate conceptual soundness, stability and explainability.
AI and machine-learning inventories and tiering
Most implementation failures begin with an incomplete inventory. AI capability rarely enters a bank through the model development function. It arrives through vendor upgrades, cloud services, productivity tooling and business-unit experimentation, none of which is routinely routed to model risk management.
- Define the model boundary in writing, including deterministic rule engines, embedded vendor analytics and generative systems used to produce or summarise decision-relevant content.
- Record for every entry the owner, the business use, the decision it influences, the data sources, the training and retraining cadence, and whether outputs are consumed automatically or by a human.
- Capture dependencies, since model chains propagate error: an AI-derived feature used inside a credit model makes the credit model’s tier the effective floor for the feeder.
- Run periodic discovery sweeps across procurement, cloud spend and change records rather than relying on self-declaration.
Tiering should be driven by consequence rather than technique. A simple model that determines customer outcomes at scale outranks a complex model used for internal exploration. Criteria that hold up under supervisory challenge include financial exposure influenced, regulatory relevance, customer impact, reversibility of the decision, degree of automation and availability of a credible alternative.
| Tier | Typical characteristics | Expected treatment |
|---|---|---|
| Tier 1 | Automated or minimally supervised decisions on customers or capital; material financial or regulatory consequence | Full independent validation before use, annual revalidation, formal monitoring with thresholds and escalation |
| Tier 2 | Human-reviewed recommendations; moderate exposure; established fallback available | Risk-based validation, periodic review, monitoring of drift and override rates |
| Tier 3 | Internal analytical or productivity use with no direct external decision | Registration, ownership, use restrictions and lightweight review |
Explainability and limitation statements
Explainability is not a single property. Supervisors distinguish between global explanation of how a model behaves across the population, local explanation of why an individual decision was reached, and the ability to justify the model’s use at all. A firm can satisfy the first and fail the second, which is precisely the position that causes difficulty in complaints, redress and conduct reviews.
- Attribution methods approximate the model rather than describing it, so their own limitations must be documented alongside their output.
- Explanations should be tested for stability: if the explanation for a decision changes materially under trivial input perturbation, it is not usable as a justification.
- Every material AI model should carry a written limitation statement describing the population it was fitted to, the conditions under which its conclusions cease to hold, and the uses that are prohibited.
- Where explainability cannot reach the required standard for a use case, the correct response is a use restriction, not a longer methodology annex.
Data quality and model drift
Machine-learning performance depends on the continued resemblance between production data and training data. Two distinct failures need separate monitoring. Data drift is a change in the input distribution; concept drift is a change in the relationship between inputs and the outcome. A model can display stable inputs while its underlying relationship decays, which is why input monitoring alone is insufficient.
- Establish lineage from source system to model input, with controls over the transformations applied in between.
- Monitor input distributions, missingness, category emergence and feature correlation structure against the training baseline.
- Monitor realised outcomes where feedback exists, and use proxy or challenger comparison where outcomes are delayed, as they are in credit.
- Set quantitative thresholds in advance with defined escalation, so that a breach triggers action rather than discussion.
- Treat retraining as a change event subject to governance, since silent retraining defeats the approval on which use permission rests.
Firms operating IFRS 9 expected-credit-loss models already run analogous monitoring for staging and forward-looking assumptions. Extending that discipline to AI components is usually more efficient than constructing a parallel regime.
Vendor and third-party AI models
Responsibility for a model outcome is not transferred by procurement. Where a vendor supplies an AI system that influences a regulated decision, the firm remains accountable under SS1/23 and must be able to evidence proportionate validation despite limited access to the underlying method.
- Negotiate documentation, testing access and change notification at contract stage; these are rarely obtainable afterwards.
- Validate on the firm’s own portfolio, since vendor performance evidence reflects the development population, not the user’s.
- Test behaviour empirically where the internal method is not disclosed: benchmark against alternatives, examine sensitivity to input perturbation and analyse outcomes across relevant segments.
- Record substitution and exit options, because concentration in a small number of AI providers is itself a resilience exposure.
- Apply the same expectations to models embedded within broader software, which are frequently missed entirely.
Human oversight and overrides
Human oversight is a control only if it is capable of changing an outcome. Oversight that consists of confirming a recommendation the reviewer cannot interrogate provides governance comfort without risk mitigation, and supervisors examine this closely.
- Give reviewers the information, time and authority to disagree, and record what they were shown at the point of decision.
- Monitor override rates in both directions: a rate near zero suggests automation bias, while a persistently high rate suggests the model is not fit for the use.
- Analyse override reasons as a model performance signal, since reviewers often detect deterioration before monitoring statistics do.
- Define escalation for cases the model is known to handle poorly, using the documented limitation statement as the trigger.
Independent validation of AI systems
The fourth principle requires validation independent of development, risk-based in intensity and empowered to constrain use. For AI systems the core validation areas are unchanged — conceptual soundness, data, implementation, performance, sensitivity and benchmarking, governance and use — but each carries additional tests.
- Conceptual soundness: is a flexible learner justified for this problem, and is the additional performance sufficient to warrant the loss of transparency relative to a simpler alternative?
- Data: is the training population representative of current business, and do proxy variables reintroduce characteristics that must not drive the decision?
- Implementation: does the production pipeline reproduce the validated model, including feature engineering, and is versioning of code, data and parameters demonstrable?
- Performance: results out of sample and out of time, by segment rather than in aggregate, with attention to sparse but material segments.
- Sensitivity and benchmarking: stability under perturbation and re-estimation, and comparison against a transparent challenger, which remains the most informative single test.
- Governance and use: confirmation that deployment matches the approved scope and that monitoring is live rather than designed.
The validation output should be a conditional opinion: fit for a stated purpose, under stated conditions, with findings that carry owners and deadlines. Unconditional approval of an AI model is rarely defensible.
DORA, cyber risk and testing
The EBA’s June 2026 Risk Assessment Report makes the operational-resilience dimension explicit: requirements under the Digital Operational Resilience Act apply to banks’ use of AI systems, and frontier AI can amplify cyber risk. This has organisational consequences. An AI model that is validated as statistically sound but hosted on infrastructure outside the ICT risk register is not adequately governed.
- Register AI systems and their supporting services within ICT risk management and third-party registers, not only within the model inventory.
- Include AI-dependent processes in resilience testing, incident classification and reporting arrangements.
- Extend security testing, including penetration testing where proportionate, to model-serving infrastructure, data pipelines and interfaces through which inputs can be manipulated.
- Consider adversarial exposures specific to models, such as input manipulation, prompt injection into generative components and extraction of training data.
- Ensure that continuity arrangements specify what the business does when an AI service is unavailable or degraded, which is the same question as fallback for a failed model.
Management-body accountability
SS1/23 places ownership of the model risk framework with the board and senior management. In practice this means an identified senior individual is accountable for the AI model estate, receives reporting sufficient to exercise that accountability, and can be examined on it.
- Board-level reporting should show the inventory by tier, validation coverage against plan, open findings by age and severity, monitoring breaches and material use restrictions.
- Risk appetite should state where AI is permitted, where it is prohibited and what degree of automation is acceptable for each decision class.
- Escalation routes must allow validation to constrain or suspend use without requiring the agreement of the business that owns the model.
- Given the PRA’s stated intention to discuss AI and machine-learning models with chief risk officers as it concludes firm reviews by the end of 2026, the chief risk officer should be able to describe the estate, its risks and its controls without recourse to a prepared script.
Practical implementation priorities for 2026
- Complete and reconcile the AI and machine-learning inventory through independent discovery, not self-declaration.
- Document the tiering methodology and evidence its consistent application, including the reasoning for borderline cases.
- Produce limitation statements and use restrictions for every tier 1 and tier 2 AI model.
- Stand up drift and outcome monitoring with pre-agreed thresholds, owners and escalation, and evidence that breaches were acted upon.
- Close vendor documentation gaps and validate purchased models on the firm’s own data.
- Align the model inventory with ICT and third-party registers so that DORA obligations and model risk obligations reference the same population.
- Rehearse the supervisory conversation: inventory, tiering rationale, validation coverage, findings, mitigants and board reporting.
Firms with a mature model risk framework generally do not need a new one. The gap is coverage: systems that entered the organisation through routes the existing framework was never designed to intercept.
Practical example
Consider a mid-sized bank that deploys a gradient-boosted model to prioritise early-arrears collections activity. The model does not grant or refuse credit, so the business initially classifies it as an operational efficiency tool outside the model inventory.
- Identification: the inventory review finds that the model determines which customers receive forbearance outreach, which affects customer outcomes and, through cure rates, the realised loss given default feeding the IFRS 9 model.
- Classification: the dependency on a regulatory model and the customer-outcome consequence place it in tier 1 rather than tier 3.
- Development and use: the approved use is prioritisation only; automated case closure is excluded and the restriction is recorded in the model documentation.
- Independent validation: a transparent challenger reproduces most of the ranking performance, so the validation opinion is conditional on segment-level monitoring and on stability of the explanation for prioritisation decisions.
- Mitigants: override monitoring, a monthly drift report against the training baseline and an annual revalidation are imposed, with escalation to the credit risk committee on threshold breach.
The substantive risk in this example was not the algorithm. It was the classification decision that would have kept an outcome-relevant model outside the framework entirely.
Limitations and caveats
- This article summarises publicly available supervisory material and does not constitute legal, regulatory or investment advice.
- Supervisory expectations evolve; firms should read the current text of SS1/23 and any subsequent PRA and EBA publications directly.
- SS1/23 applies to banks within its stated scope; insurers and other regulated firms are subject to different, though frequently analogous, expectations.
- Tiering thresholds, monitoring metrics and validation depth are firm-specific and must be calibrated to the institution’s own estate and risk appetite.
- No proprietary methodology, architecture or calibration approach is described here.
Conclusion
AI model risk management in 2026 is an exercise in coverage and evidence rather than in new theory. The five SS1/23 principles already describe what is required; the difficulty is that artificial-intelligence systems enter banks through channels that traditional model governance does not monitor.
The firms that will withstand supervisory review are those that can produce a complete inventory, a defensible tiering rationale, independent validation opinions with conditions attached, live monitoring with evidence of action, and a chief risk officer who can discuss the estate without preparation.
With the PRA intending to conclude its reviews of remaining firms by the end of 2026 and the EBA linking AI adoption to cyber risk and DORA obligations, the practical priority is to align model risk and operational resilience around a single, complete view of the AI estate.
References
- SS1/23 – Model risk management principles for banks — Bank of England / Prudential Regulation Authority, 2026
- Prudential Regulation Authority Annual Report 2025/26 — Bank of England / Prudential Regulation Authority, 2026
- Risk Assessment Report — European Banking Authority, 2026
- Regulation (EU) 2022/2554 on digital operational resilience for the financial sector (DORA) — EUR-Lex, 2022
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence — European Commission / EUR-Lex, 2024
Author bio

Jonas Mohamed Osman Abdelghafour, known as Yonas Osman is an actuary, FRM and financial risk professional specialising in banking, insurance, model risk, capital modelling and quantitative risk management.
