FCA Expects AI Controls Testing Against Frontier Threats: What Second-Line Teams Need Now
“Not theory. Evidence.” That is how the FCA’s Head of Innovation, Colin Payne, framed what the regulator wanted when it reopened its AI Input Zone on 14 May 2026. Two sentences, four words, and a fairly complete statement of where UK supervision of AI is heading.
Second-line teams should read it as a warning about their own artefacts. A policy is theory. A control that has been tested against a realistic threat, with a dated record of the test and what it found, is evidence. Most firms have a great deal of the first and very little of the second.
Research and drafting for this article were AI-assisted. Every figure is linked inline to the source that published it. Editorial responsibility is mine.
What Changed on 15 May
The day after the Input Zone reopened, the Bank of England, the FCA and HM Treasury published a joint statement on frontier AI models and cyber resilience. Its central claim is worth stating plainly, because it is unusual for regulators to put it this bluntly: the cyber capabilities of current frontier models already exceed what a skilled human practitioner can achieve, and they do it faster, at greater scale, and for less money.
Firms were told to strengthen their defences accordingly. The statement was not framed as a consultation.
For a risk or compliance function, the operational consequence is uncomfortable. Controls designed against a threat model where the attacker is a competent human with finite time do not automatically hold when the attacker is a competent human with a model that never sleeps. And it is the second line that will be asked to demonstrate they hold.
The Trajectory This Sits On
None of this arrived from nowhere. In January 2026 the House of Commons Treasury Committee reported that the FCA, the Bank and the Treasury were not doing enough on AI risk and were exposing consumers and the financial system to potentially serious harm. It asked for comprehensive practical guidance by the end of 2026 on two specific questions: how existing consumer protection rules apply to AI, and what assurance senior managers must be able to give for harm caused through it.
That guidance is now being assembled from live submissions. The Input Zone asked market participants for concrete examples of good and poor practice, closed on 19 June, and will feed a good-and-poor-practice publication later in the year. Non-binding, and it will function as the benchmark supervisors reach for. Running alongside it, the Mills Review, launched on 27 January under Sheldon Mills, is looking at how AI reshapes retail financial services out to 2030, with four themes covering AI evolution, market impact, consumer trends and regulatory adaptation. Recommendations are expected over the summer.
Ropes & Gray reads the whole sequence as a move from principle to practice. I would put it more narrowly. The FCA has decided to build its supervisory standard out of what firms are actually doing, which means the firms that submit real evidence now are helping to write the thing they will later be measured against.
The Five Expectations, Read Closely
The joint statement sets expectations across five areas. Each converts into work for a second line, and the conversion is not always obvious.
Governance and strategic direction. Boards and senior management need sufficient understanding of frontier AI risk to set strategy and oversee control functions. Read carefully, that is not an awareness expectation. It requires a board to articulate how these capabilities change its own firm’s risk profile, and what it decided to spend as a result.
Vulnerability management at scale. Triage, prioritise and remediate more quickly, more frequently and at scale, with automation where it helps. The statement is explicit that frontier models can find zero-day vulnerabilities across an enterprise estate. A quarterly assessment cycle is not a slow version of the right answer. It is the wrong unit of time.
Third-party risk. External applications, libraries and services inside the firm’s network must be monitored for vulnerabilities, open-source components included. For any firm whose AI pipeline leans on third-party models, inference providers and a long tail of packages, that is a direct hit.
Defensive capability. Firms should consider automated and AI-enabled defences that operate at a speed comparable to AI-driven attacks. The implication is that manual review cannot keep pace, which puts every periodic control in the second line under question.
Incident response. Rapid response and recovery aligned to the effective practices the Bank published in October 2025. Tested playbooks, not documented ones. The distinction is the whole point.
Who Is Personally Accountable
Addleshaw Goddard asked the question directly in January 2026: can a senior manager be personally liable under the senior managers regime for decisions an AI system made inside their business?
They can. David Geale, the FCA’s Executive Director for Payments and Digital Finance, told the Treasury Committee that the regime keeps individuals “on the hook” when consumers are harmed through AI. Handing a decision to a model does not move the liability anywhere.
The complication is that there is no dedicated senior manager function for AI. Accountability is distributed, usually across chief operations, chief risk and the relevant business heads, which creates an attribution problem the moment something goes wrong. Addleshaw Goddard frames the hard case well. A manager approves a tool that handles 95% of cases correctly and causes harm in the remainder, having put governance and mitigation in place. Were the steps taken reasonable? Their answer, in the current state of the regime, is that it is arguable.
Arguable is not a position any named individual should be content to occupy.
That ambiguity will not survive the FCA’s practical guidance, and second-line teams should not wait for it to be resolved externally. Map each AI deployment to a named function holder. Record the control reliance at the moment of sign-off, in the person’s own words rather than the project team’s. Define internally what reasonable steps look like for this class of system, and do it before somebody else defines it for you under examination conditions.
The Distance Between Expectation and Evidence
The gap between what supervisors now expect and what most firms could produce on request is wide, and it is measurable.
The Cambridge Centre for Alternative Finance surveyed 628 organisations across 151 jurisdictions for its 2026 study and found 81% of financial services firms adopting AI, against 48% of the 130 regulatory authorities surveyed still describing themselves as exploring it. Governance is behind deployment on both sides of the supervisory relationship.
McKinsey’s 2026 trust work puts only 30% of organisations at maturity level three or above in governance and controls, roughly two levels behind their own deployment ambitions. And Grant Thornton’s 2026 survey found 78% of executives without strong confidence that they would come through an independent audit of AI governance in 90 days.
Set that last figure beside the good-and-poor-practice publication and the significance changes. That document will, in effect, describe what an independent look at AI governance consists of. Firms that cannot show functioning controls, tested and dated, will be measured against a definition they had no part in shaping.
Working back from the FCA’s published positions and Aveni’s mapping of them, the expectations are crystallising around five things a firm should be able to produce: a pre-deployment risk assessment tied to Consumer Duty outcomes, named accountability with the control reliance documented, real-time monitoring with defined intervention thresholds, interaction-level audit trails covering what the system did and why, and third-party resilience evidence from audit rights that were actually exercised.
Where a Second Line Should Spend the Next Two Quarters
The period between now and the guidance is not a waiting room. It is the only preparation window that exists.
Tie every AI deployment to a Consumer Duty outcome. Any system touching a customer-facing process or shaping a customer result should be assessed against products and services, price and value, consumer understanding and consumer support. The Duty is the lens the FCA has said it will use, so the assessment should be structured the way the FCA structures the Duty, not the way the project structured the build.
Get accountability in writing. For each system, identify the function holder in whose area it sits, and record their own assessment of the controls at deployment. This is not a governance formality. It is the document that gets read first when a poor outcome surfaces.
Test controls rather than describing them. Guidehouse’s 2026 framework sets out a nine-step control lifecycle, and the parts most firms skip are codified prompt tests, red-team exercises that expose abuse paths, drift detection and adversarial verification of guardrails. Commission those. Do not let internal audit be the function that discovers they were never run.
Move monitoring from periodic to continuous. Thresholds that trigger intervention before a customer is affected, with intervention logs kept and reviewed. Sampling two or three per cent of AI-driven interactions will not surface a systematic failure running through thousands of decisions a day, and a supervisor will say so.
Use the audit rights you negotiated. Model providers, cloud infrastructure and inference services all fall inside the third-party expectations. A signed contract evidences a contract. Exercised audit rights, tested incident procedures and verified exit arrangements evidence resilience.
Somebody Is Going to Define What Good Looks Like
The FCA has told the market how it intends to set the standard: out of submitted examples of what firms actually do, published as good and poor practice, then used in supervision.
That is an unusual amount of influence to hand to the regulated population, and it is time-limited. The firms with tested controls and dated records will find their practice reflected in the benchmark. The firms that submitted nothing will be measured against somebody else’s operating model, with a named individual answering for the difference.
Evidence, not theory. They were fairly clear about it.
Regulatory position as at July 2026. The FCA's good-and-poor-practice publication and the Mills Review recommendations were both still outstanding at that date.
Primary Documents and Further Reading
- Bank of England, FCA and HM Treasury, Joint statement on frontier AI models and cyber resilience, 15 May 2026. https://www.bankofengland.co.uk/news/2026/may/boe-fca-and-hm-treasury-joint-statement-on-frontier-ai-models-and-cyber-resilience
- FCA, FCA, Bank of England and Treasury joint statement on frontier AI models and cyber resilience. https://www.fca.org.uk/news/statements/fca-boe-treasury-joint-statement-frontier-ai-models-cyber-resilience
- House of Commons Treasury Committee, AI in Financial Services, 20 January 2026. https://committees.parliament.uk/publications/51128/documents/283671/default/
- Regulation Tomorrow, FCA AI Input Zone re-opens, 14 May 2026. https://www.regulationtomorrow.com/2026/05/fca-ai-input-zone-re-opens/
- FCA, Review into the long-term impact of AI in retail financial services (the Mills Review), 27 January 2026. https://www.fca.org.uk/publications/calls-input/review-long-term-impact-ai-retail-financial-services-mills-review
- Freshfields, The FCA looks to 2030: key takeaways from the Mills Review, February 2026. https://www.freshfields.com/en/our-thinking/briefings/2026/02/the-fca-looks-to-2030-key-takeaways-from-the-mills-review-on-ai-in-retail-financial-services
- Ropes & Gray, From Principles to Practice: the FCA’s evolving expectations on AI governance, June 2026. https://www.ropesgray.com/en/insights/viewpoints/2026/06/102n7e3/from-principles-to-practice-the-fcas-evolving-expectations-on-ai-governance
- Addleshaw Goddard, Can senior managers be liable for decisions made by AI?, 2026. https://www.addleshawgoddard.com/en/insights/insights-briefings/2026/global-investigations/senior-managers-liable-under-uk-regulatory-regime-decisions-made-ai/
- Cambridge Centre for Alternative Finance, 2026 Global AI in Financial Services Report. https://www.jbs.cam.ac.uk/faculty-research/centres/alternative-finance/publications/2026-global-ai-in-financial-services-report/
- McKinsey, State of AI trust in 2026: shifting to the agentic era. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/tech-forward/state-of-ai-trust-in-2026-shifting-to-the-agentic-era
- Grant Thornton, 2026 AI Impact Survey. https://www.grantthornton.com/services/advisory-services/artificial-intelligence/2026-ai-impact-survey
- Guidehouse, Operationalising AI governance: risk management, controls and testing, 2026. https://guidehouse.com/insights/financial-services/2026/operationalizing-ai-governance
- Aveni, How to evidence AI agent compliance: five expectations from the FCA. https://aveni.ai/blog/evidence-ai-agent-compliance-fca/
- Aveni, Consumer Duty four outcomes: what the FCA expects you to evidence in 2026. https://aveni.ai/blog/consumer-duty-outcomes/
- Global Policy Watch, UK financial services regulators’ approach to AI in 2026, April 2026. https://www.globalpolicywatch.com/2026/04/uk-financial-services-regulators-approach-to-artificial-intelligence-in-2026/