← Back to all blogs

Louise Cermak | 26 August 2026

Your AI Programme Will Find Your Data Problems. Better to Find Them First.

Public sector organisations do not need AI to create data problems. Most already have them.

Duplicate records. Uncontrolled spreadsheets. Legacy and unmanaged SharePoint sites. Overly broad or outdated permissions. Documents retained long beyond their review date. Critical knowledge buried in PDFs, emails and shared drives.

The story is familiar.

People have learned to work around these weaknesses. They know which spreadsheet to trust, who owns a particular dataset and which version of a document is current. Manual access creates friction too as someone has to know the information exists, find it, open it and interpret it.

AI changes that equation.

Connect retrieval tools, co-pilots or agents to the same estate and information can be found, combined and reused far faster and at much greater scale. The underlying weaknesses may not have changed but their reach and consequences have.

That is why AI data readiness is becoming a practical leadership issue, not simply a data-quality exercise.

It is also why data foundations are high on the agenda at this year’s Public Sector Data & AI Summit, taking place on 2 September at One Great George Street, Westminster. Catapult CX will join public sector data, digital and AI leaders to discuss what it takes to move AI beyond experimentation and into trusted, operational use.

Download the AI Readiness Playbook

AI exposes existing data problems and amplifies their impact

The UK Government’s 2026 guidance on making datasets ready for AI makes an important distinction. AI-ready data is not defined by format alone. It depends on context, governance, interoperability and whether the data is suitable for the specific AI use case.

That matters because data can work adequately for its original purpose but become unreliable, unsafe or misleading when reused by AI.

Consider what changes when the friction of manual access disappears.

Existing data weakness Why it may remain manageable today What changes with AI
Duplicate or conflicting records Experienced staff know which source to trust AI may retrieve different versions and produce inconsistent answers
Uncontrolled spreadsheets Their use is often local and limited Once connected to AI workflows, locally maintained information can influence activity across the organisation
Overly broad or outdated permissions Exposure may depend on someone knowing where the information is stored AI can make technically accessible information far easier to discover, combine and reuse
Retained or obsolete content Old material can remain buried and rarely accessed AI can rediscover and reuse information that people had effectively forgotten
Poorly structured documents People can apply judgement and context when interpreting them Without sufficient metadata, provenance and context, AI may struggle to determine what information means and whether it should be trusted

The point is not that AI automatically makes every one of these problems dangerous. It is that organisations need to understand which weaknesses could become material before connecting powerful retrieval, decision-making and automation capabilities to their estate.

The risk increases further with agentic AI. Agents may have access not only to information, but also to the systems and tools that allow them to act on it. When those actions happen faster than people can meaningfully review them, least-privilege access, tightly bounded use cases and effective human oversight become essential.

The public sector’s ‘digital heap’ is becoming part of the AI estate

This problem is particularly relevant to government because so much valuable information is unstructured.

Documents, emails, Teams conversations, SharePoint sites and legacy shared drives have accumulated around human workflows, record-keeping and collaboration – not machine retrieval.

Government guidance describes the accumulated content across these repositories as the ‘digital heap’. Some of it is valuable. Some is duplicated, obsolete or trivial. Some carries retention requirements. Some lacks the context needed to determine what should happen to it.

AI creates an opportunity to make this information more accessible and useful. But it can also surface and reuse weaknesses that were previously difficult to find.

That is why the right question is not simply; Do we have enough data for AI?

It is; Can the data required for this use case be found, accessed appropriately, interpreted correctly and reused with confidence?

That is a much more useful definition of AI data readiness.

Do not try to clean the entire estate before starting

The answer is not a three-year enterprise data-cleaning programme.

Government guidance emphasises that AI readiness should be assessed against specific datasets and specific use cases. Different applications place different demands on data.

A knowledge assistant answering internal policy questions needs authoritative documents, appropriate permissions, clear provenance and enough metadata to distinguish current policy from obsolete guidance.

A predictive service may need structured historical data, consistent definitions and evidence that the dataset is suitable for the decision being supported.

An agent capable of taking action requires another level of control because access to data is connected to operational authority.

So start with the use case.

What outcome are you trying to improve? What data does it depend on? Where does that data live? Who owns it? Can the right systems access it? Which weaknesses could materially affect accuracy, security, legality or trust?

Then decide what actually needs fixing.

That is a much smaller and more defensible problem than trying to make the entire data estate AI-ready.

Download the AI Readiness Playbook

Data readiness cannot be assessed in isolation

Data may be central to AI readiness, but it is only one part of the environment in which AI must operate.

A dataset may be findable, accessible and well governed, but the use case can still fail if legacy systems cannot provide reliable access, the architecture cannot support integration, security controls are inadequate or the organisation lacks the governance needed to move from experimentation into production.

Weak foundations do more than increase risk. They increase costs, slow delivery and make it harder for AI investment to produce meaningful operational outcomes.

That is why Catapult CX assesses data as part of the wider technology and operating environment, not as a separate workstream.

Diagnose your readiness for AI

Catapult CX starts with the outcome the organisation is trying to achieve. What operational problem needs solving? Which systems and workflows are involved? Where could AI realistically replace, augment or improve the way the organisation works? And is AI genuinely the right response?

Catapult CX’s Data and AI Readiness Diagnostic then assesses the foundations supporting each potential use case across six connected areas:

  • Data
  • Technology
  • Architecture
  • Governance
  • Security and compliance
  • AI

Within the data assessment, the FAIR principles of Findability, Accessibility, Interoperability and Reusability, help determine whether the required information can be found, accessed appropriately, combined with other data and reused with confidence.

Catapult CX’s AI Readiness Playbook and Scorecard support the wider assessment, providing a structured view of the conditions that will enable or constrain delivery.

This connected approach reveals more than whether the data is ready. It shows how the data, systems, architecture, controls and operating model interact and where weaknesses in one area could prevent the use case from succeeding.

Catapult CX’s AI and Data Readiness services can be purchased through Technology Services 4 and G-Cloud 15, providing public sector organisations with established routes to engage.

Prioritise what matters most

Most organisations have more potential AI use cases than they can pursue at once. They also have more weaknesses in their estate than they can, or need to, fix immediately.

The purpose of the diagnostic is therefore not to produce an indiscriminate list of problems. It is to determine which opportunities matter, which constraints are material and where investment should be focused first.

Catapult CX prioritises use cases and remediation according to their operational value and public service impact, readiness to deliver, risk and governance requirements, technical dependencies, likely cost and complexity, the consequences of leaving weaknesses unresolved and whether the opportunity can be tested safely within a bounded use case.

The output is not simply a collection of maturity scores. It is a practical gap register and prioritised roadmap showing:

  • which AI opportunities can move forward;
  • which foundations need strengthening first;
  • the dependencies that could affect delivery;
  • what can be tested in the near term;
  • what should wait;
  • how activity should be sequenced around the outcomes that matter.

This gives leaders a defensible basis for deciding where to invest, what to address and what not to do.

Modernise iteratively

Readiness should lead to action, but not necessarily to a large transformation programme.

Catapult CX then helps organisations modernise iteratively around their priority use cases. That may involve improving the quality, accessibility or governance of specific data; modernising the systems and integrations on which the use case depends; strengthening security and compliance controls; or creating a more coherent architecture for AI to operate within.

The aim is to address the foundations that materially affect delivery, test the use case in controlled increments and use the evidence generated to shape the next stage.

This avoids two common extremes:

  1. Trying to fix the entire estate before beginning
  2. Or layering AI over systems, data and processes that are not ready to support it.

The same principle runs through Catapult CX’s wider modernisation work. For the UK Ship Register, Catapult CX addressed the underlying systems, data and operational processes rather than adding a digital interface to a broken operating model. Following the transformation:

  • data and certification errors fell by 89%
  • monthly transactions increased by 625%
  • cost per transaction fell by 45%

It was not a specific AI programme, but it demonstrates what strong, connected foundations make possible – lower costs, faster delivery and better operational outcomes.

AI makes those foundations more important, not less.

Diagnose your readiness for AI. Prioritise what matters most. Modernise iteratively.

Find the problems before AI does

The Public Sector Data & AI Summit on 2 September 2026 is focused on the same challenge – moving AI beyond experimentation and into trusted, governed and operational use at scale.

Catapult CX will be at One Great George Street, Westminster, discussing how public sector organisations can diagnose their readiness for AI, prioritise what matters most and modernise. If you are determining where AI can create value across your organisation and whether your data, technology and operating environment are ready to support it, come and talk to us.

The question is not whether your organisation has data problems. Most do.

The question is which of those problems will become material when AI can find, combine and act on information faster than your existing operating model was designed to handle.

Find those first. Then build.

Download the AI Readiness Playbook