← Back to all blogs

Louise Cermak | 27 July 2026

Why Enterprise AI Fails After the Pilot. The Post-PoC Gap in Regulated Financial Services

AI in Public Sector

Most enterprise AI pilots do exactly what they are designed to do. They prove the model works. What they rarely prove is whether the organisation is ready to operate that model as part of a live, regulated business. In UK financial services, AI pilots typically succeed technically. They stall because a successful proof of concept never answers the operational, governance and resilience questions that arise once the system moves into production.

The moment an AI pilot begins supporting live business processes, it stops being solely a technology initiative. It becomes an operational resilience, compliance and model risk challenge. The FCA’s September 2025 feedback statement, FS25/5, reinforces that this transition is now recognised at regulatory level, not simply as a delivery challenge for engineering teams.

This is not a reason to slow AI adoption. It is a reason to be more precise about what production-ready actually means in a regulated environment. Enterprise AI programmes rarely fail because the technology is incapable. They fail because organisations have not defined the governance, ownership and operational controls required to run that technology safely at scale.

Take the AI Readiness Scorecard

This isn’t the generic ‘pilot-to-production’ problem

There is a familiar version of this story. AI pilots are built on clean, curated data with a small, motivated team and a clearly defined objective. Production environments are more complex. Data changes, integrations multiply, ownership becomes shared and operational realities replace controlled testing. We have written previously about why AI initiatives struggle to move from pilot to production with operating model gaps, unclear business and technical ownership, and the assumption that a successful demonstration is the same as production readiness.

Regulated financial services introduces another layer entirely. When an AI pilot begins supporting a customer journey, credit decision, claims process or compliance workflow, it may fall within the scope of the FCA’s operational resilience expectations and the PRA’s model risk management framework. Those governance, resilience and accountability requirements do not apply in the same way to an isolated proof of concept or sandbox environment.

That is why many technically successful AI pilots never become production systems. The challenge is no longer proving the model works. It is proving the organisation is ready to operate it safely, consistently and under regulatory scrutiny.

What ‘proof of concept paralysis’ actually looks like

Proof of concept paralysis is an industry term rather than an FCA definition. It describes the point where an AI pilot succeeds technically but never progresses into production because the operational, governance and regulatory questions remain unresolved.

The pattern is reflected throughout the FCA’s FS25/5 feedback statement. Respondents to the AI Live Testing engagement paper noted that AI models are rarely ready to use ‘out of the box’. They require bespoke training data, integration with production systems and a much clearer understanding of how they will operate in a live environment.

In practice, the symptoms are familiar. The model works. Accuracy is acceptable. Internal stakeholders are impressed. Then progress slows because no one has defined ownership, monitoring, fallback procedures, vendor accountability, data lineage or governance once the system goes live. The pilot is not the problem. The absence of a production-readiness plan is.

The table below, highlights some of the most common reasons technically successful AI pilots fail to become governed production systems.

Pilot symptom Regulatory risk Production-readiness question
Accuracy looks good in a controlled test Operational readiness has not been demonstrated Has the model been validated against live process conditions, edge cases and failure scenarios?
No clear production owner Governance and accountability remain unclear Who owns the outcome, the controls and the decision to keep the system live?
Vendor model works in sandbox Third-party operational resilience has not been assessed What evidence demonstrates that the supplier, data flows and fallback arrangements can support a regulated service?
No drift or performance monitoring Model risk controls are incomplete How will performance, bias, drift and exceptions be monitored after deployment?
Pilot data was manually prepared Data lineage and operational repeatability have not been proven Can the production data pipeline operate reliably without manual workarounds?

If your organisation is still evaluating AI opportunities, our AI Readiness Assessment focuses on the questions that should be answered before a pilot begins. This article assumes the pilot already exists and examines the next challenge – moving from a successful proof of concept to a governed production deployment. For organisations also assessing data security and deployment models, see our guide to Private AI for Financial Services.

The FCA’s evidence. What FS25/5 reveals

The FCA’s FS25/5 feedback statement is significant because it confirms that the challenge of moving AI from proof of concept into live operation is no longer just an industry concern. It is now recognised by the regulator as a practical governance and operational readiness issue.

Published on 9 September 2025, FS25/5: AI Live Testing summarises responses to the FCA’s April 2025 engagement paper. AI Live Testing, delivered through the FCA’s AI Lab, gives selected firms a structured environment to test AI systems with direct regulatory engagement before wider deployment.

The programme is voluntary and aimed at firms that already have a working AI proof of concept operating within UK financial markets. The first cohort began in October 2025, the second application window ran from 19 January to 24 March 2026, and the FCA expects to publish its evaluation findings in Q1 2027.

However, you do not need to participate in AI Live Testing for FS25/5 to be relevant. The feedback reflects issues the FCA is already seeing across the market – organisations with technically successful AI pilots that have not yet addressed the governance, operational resilience and accountability required to operate those systems in production.

It is equally important to understand what FS25/5 does not introduce. The FCA is not creating a standalone AI rulebook. Instead, as Freshfields observes, the direction of travel is to govern AI through existing regulatory frameworks, including the Senior Managers and Certification Regime, operational resilience requirements and model risk management, rather than introducing an entirely separate AI regulatory regime.

The operational resilience trap. PS21/3 and important business services

This is where the gap between a successful AI pilot and a production-ready system becomes tangible. Under PS21/3: Building operational resilience, in-scope firms must identify their important business services – services that, if disrupted, could cause intolerable harm to customers or threaten the stability of the UK financial system.

For each important business service, firms must define impact tolerances, map the people, processes, technology and third parties that support it, and demonstrate that the service can remain within those tolerances under severe but plausible disruption scenarios. Full compliance has been required since 31 March 2025.

PS21/3 was not written specifically for AI, but it does not need to be. The moment an AI system supports an important business service, whether a credit decision engine, claims triage process or KYC workflow, it becomes part of that service’s operational resilience obligations. A model that has not been mapped, scenario-tested, monitored and assigned appropriate fallback arrangements is not production-ready, regardless of how well it performed during the pilot.

This is one of the reasons AI programmes often stall after a successful proof of concept. The challenge is no longer the model itself. It is demonstrating that the wider service can continue to operate safely, consistently and within regulatory expectations when AI becomes part of day-to-day operations.

For firms whose AI pilots depend on legacy platforms or fragmented data estates, that mapping exercise is often more difficult than expected. Legacy modernisation therefore becomes an enabling step for AI adoption rather than a separate technology programme.

The model risk gap. SS1/23 and what changes once AI scales

Alongside operational resilience sits model risk.

The PRA’s Supervisory Statement SS1/23, Model Risk Management Principles for Banks, is technology-agnostic. It was not written exclusively for AI, but it was designed to capture the governance and risk management principles that apply equally to AI and machine learning models.

The PRA reinforced that position. In November 2025, following discussions with 21 PRA-regulated firms, it published a summary of its model risk roundtable on AI. The focus was practical rather than theoretical, namely, how the principles in SS1/23 should be applied as AI systems move into live business operations.

This is another point where the gap between a successful pilot and a production-ready system becomes clear. Once an AI model begins influencing live business decisions, the expectations change. Validation, continuous monitoring, drift detection, audit trails, change control and governance become operational requirements rather than optional good practice.

This is not compliance for compliance’s sake. A model that influences customer outcomes, financial decisions or regulatory processes needs demonstrable oversight throughout its lifecycle in a way that an isolated sandbox pilot never has to prove.


Take the AI Readiness Scorecard

Third-party and outsourcing risk at scale

Third-party AI can accelerate adoption, but it also makes the transition from pilot to production more complex. In the Bank of England and FCA’s 2024 survey of AI and machine learning in UK financial services, third-party implementations accounted for around one-third of all AI use cases, up from 17% in the 2022 survey.

Greater reliance on third-party models does not reduce accountability. If a vendor-supplied model supports an important business service, the regulated firm remains responsible for understanding the dependency, the data flows, the operational resilience implications and the fallback arrangements. Accountability for customer outcomes and regulatory compliance cannot be outsourced.

This is another reason AI pilots that perform well in testing can struggle in production. A model may deliver acceptable technical performance, but unless the supplier relationship, governance controls, operational dependencies and resilience measures have all been mapped and validated, the organisation has not yet demonstrated production readiness.

What production readiness looks like

Organisations that successfully move AI from pilot to production treat governance as part of the platform, not as a compliance exercise that begins after deployment. Monitoring, ownership, resilience and operational controls are designed into the solution from the outset, ensuring the model can be managed safely throughout its lifecycle rather than simply demonstrating good performance on day one.

This is where MLOps becomes operationally important. Monitoring, drift detection, model versioning, release management and rollback are built into the deployment process rather than added after go-live. Production readiness is not achieved when a model reaches an acceptable level of accuracy. It is achieved when the organisation can operate that model safely, consistently and with confidence.

Where organisations require AI agents or retrieval-based systems that integrate with their own data and infrastructure, Custom Agent & RAG Development provides greater control over model selection, hosting, data flows and governance. That becomes increasingly important where sensitive information, regulated workflows or third-party dependencies form part of the production environment.

None of this is quick or inexpensive. Mapping important business services, validating models, building monitoring pipelines, establishing governance processes and reviewing third-party dependencies all require investment. However, this is the work that transforms a technically successful pilot into a dependable production capability.

Catapult CX’s identity verification platform demonstrates the same principle in practice. Verification accuracy increased from 50% to 97%, model retraining reduced from two days to around 1.5 hours, infrastructure costs fell by more than 96%, and throughput increased by 3,100%. The significance is not the metrics alone. Monitoring, retraining and operational control were designed into the platform from the beginning, rather than treated as post-deployment maintenance.

External research reaches the same conclusion. RAND Corporation’s 2024 study found that AI projects commonly fail because organisations misunderstand the problem they are solving, lack the necessary data, underinvest in infrastructure or struggle to deploy and maintain systems over time. In regulated financial services, those delivery challenges quickly become governance, resilience and operational risk issues.

Moving beyond the pilot

A successful AI pilot is an important milestone, but it is not evidence that an organisation is ready to operate AI in production. If operational resilience, model risk, governance, ownership and third-party dependencies have not been addressed, the pilot has not failed, it has simply reached the point where a different set of capabilities is required.

The next phase is less about improving the model and more about strengthening the operating environment around it. That means establishing governance, defining ownership, validating production readiness and building the operational controls needed to support AI safely at scale.

This is where Catapult CX’s AI Advisory work focuses, helping organisations determine where AI should be used, what governance and controls are required, who owns the associated risks and what must be true before a system moves into live operation.

Once that foundation is in place, MLOps provides the operational capability to deploy, monitor, retrain and manage AI systems throughout their lifecycle. Where organisations need AI agents or retrieval-based systems built around private data, internal knowledge or regulated workflows, Custom Agent & RAG Development provides greater control over model selection, data flows, hosting and governance.

If you are still exploring AI opportunities, the AI Readiness Assessment and AI Readiness Playbook provide a structured starting point.

If you already have a successful pilot, the more important question is no longer whether the model works. It is whether your organisation is ready to operate it.


Take the AI Readiness Scorecard

FAQs

Why are AI pilots failing in enterprises?

Most AI pilots fail because organisations underestimate what changes between a controlled pilot and live production. Data quality, integration, ownership, monitoring, governance and operational resilience all become more complex at scale. RAND Corporation’s 2024 research found that, by some estimates, more than 80% of AI projects fail, roughly twice the failure rate of comparable IT projects that do not involve AI.

What is proof of concept paralysis?

Proof of concept paralysis is an industry term for an AI pilot that technically succeeds but never moves into live production because operational, governance, resilience or third-party risk questions remain unresolved. While the FCA does not use the term, FS25/5 highlights many of the challenges that contribute to this pattern.

How do you know an AI pilot is ready for production?

An AI pilot is ready for production when the organisation has addressed governance, operational resilience, ownership, monitoring, model validation, rollback, third-party dependencies and ongoing operational control, not simply when the model achieves acceptable accuracy.

What is the difference between an AI pilot and a production AI system?

An AI pilot demonstrates that a model can work under controlled conditions. A production AI system must also be governed, monitored, resilient, explainable and capable of operating reliably within live business processes and regulatory expectations.

Does the FCA have special rules for AI in production?

No. The FCA has not introduced a standalone AI rulebook. Instead, AI is expected to be governed through existing regulatory frameworks covering accountability, operational resilience, conduct, outsourcing, data protection and governance.

What is an important business service under PS21/3?

An important business service is a service that, if disrupted, could cause intolerable harm to customers or threaten the stability of the UK financial system. Firms must identify these services, set impact tolerances, map the resources that support them and test their ability to remain within those tolerances during severe but plausible disruption.

Does SS1/23 apply to AI models?

SS1/23 is technology-agnostic, so it does not specifically regulate AI. However, its model risk management principles apply equally to AI and machine learning models and the PRA has published additional guidance explaining how those principles should be interpreted in practice.

Does using a third-party AI model reduce regulatory responsibility?

No. Even where an AI model is supplied by a third party, regulated firms remain responsible for governance, operational resilience, model oversight and customer outcomes. Accountability cannot be outsourced.

What is MLOps and why does it matter?

MLOps is the set of operational practices used to deploy, monitor, maintain and improve machine learning models in production. It includes monitoring, drift detection, version control, deployment, rollback and retraining, helping organisations operate AI systems safely and reliably over time.