A COO at a London consumer-credit fintech has a working AI agent that handles payment-holiday requests end to end. It's accurate in testing, customers like it, and the business case is obvious. Then the compliance director asks one question in the go-live meeting: "Which senior manager is signing this off, and what are they signing?" Nobody in the room can answer, because the agent was built to resolve tickets, not to produce evidence. The launch slips a quarter. This is the most common way UK financial-services AI projects stall in 2026, and it isn't a model problem. The FCA has stated directly that it does not plan to introduce extra regulations for AI and will rely on existing frameworks, which means Consumer Duty and the Senior Managers and Certification Regime (SM&CR) already govern your agent today. The question is whether your architecture can prove it.
What rules actually apply to an AI agent at an FCA-regulated firm?
The same rules that apply to the human process it replaces. The FCA, PRA and Bank of England have kept a technology-neutral, principles-based stance, overseeing AI through existing regulatory frameworks rather than bespoke AI-specific rules: the Consumer Duty, SM&CR, and the PRA's Supervisory Statement SS1/23 on model risk management for banks.
That sounds reassuring until you notice what it implies. There is no AI checklist to tick. Instead, every outcome the agent produces is judged against Consumer Duty's four outcomes (products and services, price and value, consumer understanding, consumer support), and every failure lands on a named human under SM&CR. Delegating a decision to a model does not delegate the liability.
Why is 2026 the year this stops being theoretical?
Because agentic AI is moving from back office to customer-facing at the largest UK lenders, and the regulator is watching it happen in real time. The FCA's second AI Live Testing cohort, announced 21 April 2026, includes Barclays, Experian, Lloyds Banking Group (Scottish Widows), UBS and GoCardless, testing use cases from credit-score insights and KYC to agentic payments.
In parallel, the House of Commons Treasury Committee warned in January 2026 that a "wait-and-see" approach risked serious consumer harm, and the FCA launched the Mills Review into how AI is reshaping retail financial services, per Covington's Global Policy Watch summary. When the big banks' agents reach customers, supervisors will benchmark everyone else against the evidence standard those firms set.
What evidence does a supervisor expect to see?
Five things, and only one of them is a document. Synthesizing the FCA's existing frameworks as they apply to agents, Aveni's breakdown of FCA expectations lands on a list that maps cleanly to system components:
- **A pre-deployment risk assessment against the four Consumer Duty outcomes**: the risks, the mitigations, and what you will monitor and how. This is the document.
- **A named senior manager** under SM&CR who has taken reasonable steps to ensure the agent is controlled effectively, and can show what those steps were.
- **Real-time monitoring with the ability to intervene** before a poor outcome reaches the customer, not a quarterly sample review after the fact.
- **An interaction-level audit trail**: what the agent did, why, and whether the outcome was good, for every conversation, not a sample.
- **Third-party and operational-resilience evidence** for any vendor model: audit rights, incident reporting, data access and an exit plan.
What does an evidence-grade agent architecture look like?
It treats every agent turn as a regulated record, not a chat log. Concretely, the architecture we build for regulated clients has five layers, each producing an artifact a supervisor can inspect:
- **Decision trace store.** Every turn writes an immutable record: the user input, the retrieved documents (with version IDs), the model and prompt version, every tool call with arguments and results, and the final response. Append-only storage with a hash chain, so the record can't be quietly edited after a complaint lands. This is the same discipline we covered in observability for AI agents, raised to evidential standard.
- **Policy-as-code action gates.** The agent can propose a payment holiday, a fee waiver or a forbearance arrangement, but a deterministic rules layer, not the LLM, decides whether that action executes or routes to a human. High-impact actions sit in an approval tier, as in our human-in-the-loop risk-tiering guide.
- **Vulnerability signal detection.** A lightweight classifier runs on every inbound message looking for the characteristics the FCA's vulnerable-customer guidance cares about: bereavement, illness, financial distress, confusion. A positive signal changes the agent's permitted behavior (no sales prompts, slower pacing, mandatory human offer) and is logged as such.
- **Outcome monitoring, not just accuracy monitoring.** Track Consumer Duty-shaped metrics per cohort: resolution without re-contact, complaint rate after agent contact, drop-off before a hardship option was offered, and outcome parity between vulnerable and non-vulnerable customers. Accuracy against a test set tells you the model works; outcome data tells the regulator customers were treated well.
- **A kill switch with a named owner.** A feature flag that drops the agent to human-only routing in seconds, owned by the accountable senior manager, with documented trigger thresholds. If nobody can turn it off, nobody can credibly claim it's controlled.
What exactly is the senior manager signing off?
Not the model, the control system around it. Under SM&CR the accountable executive must show reasonable steps, so the sign-off pack should be the risk assessment, the failure-mode inventory (what happens on hallucination, tool failure, vulnerable customer, out-of-scope request), the monitoring thresholds that trigger the kill switch, and the residual-risk statement they are accepting.
That reframing matters commercially. A senior manager asked to approve "an AI agent" will rationally say no. A senior manager asked to approve a bounded system with tested failure modes, live outcome dashboards and a switch they personally control can say yes, and that is usually the difference between a pilot and production. If the agent also makes solely automated decisions with significant effects, layer in the UK GDPR/DUAA safeguards from our UK automated decision-making guide.
How does this compare with EU and US expectations?
The UK is principles-based, the EU is rules-based, and the evidence stack is nearly identical. A firm also selling into the EU will face the AI Act's high-risk obligations for credit-scoring systems, whose deadline moved to December 2027, as we covered in the EU AI Act financial-services delay. US lenders face fair-lending and UDAAP scrutiny of the same outcomes.
The practical upshot: build the decision trace, outcome monitoring and human-override layers once, and they satisfy the FCA's principles, the EU's logging and human-oversight articles, and a US examiner's request for adverse-action reasoning. Retrofitting each jurisdiction separately is where compliance budgets go to die.
How AIBOOTSTRAPPER helps
Audit-grade traceability is the part of AI compliance work we've already shipped. For ComplyNexus, a Hong Kong compliance platform (see case studies), we built a RAG-powered engine that maps regulatory changes to a client's control library with full audit trails: 92% less manual review time, regulatory change turnaround cut from 3 weeks to 2 hours, and 100% audit-ready traceability. The same architecture pattern, every AI decision linked to its sources, versions and approvals, is what an FCA supervisor needs to see behind a customer-facing agent.
If you're a UK fintech, lender or wealth platform with an agent stuck in the go-live meeting, we design the evidence layer alongside the agent so the senior manager has something real to sign. See our AI automation and product services or book a call.
Want this done for you?
Book a free strategy call and we'll show you how to build and market your business with AI.
