Data Readiness
Four statements
The question the audit does not ask
Every enterprise AI programme runs a data readiness exercise. Almost all of them run the wrong one. The exercise that gets run is a data quality audit. Completeness, accuracy, duplication, lineage, governance sign-off. It is usually owned by the data function, it produces a scorecard, and the scorecard usually comes back acceptable. The programme proceeds on the strength of it.
Then the agent goes live and cannot tell a customer whether their broadband will be working on Friday.
Nothing in the audit was wrong. The audit answered the question it was asked. A data quality audit asks whether the data is correct. A data readiness assessment asks whether the data can be used to make a specific decision, for a specific customer, inside a specific time budget, with a defensible basis for the answer. Those are not variations of the same question, and only one of them predicts whether AI will perform in a customer conversation.
A readiness assessment asks four things a quality scorecard never records. Can this be retrieved while the customer is still on the line. Is it true at the moment the answer is given rather than at the moment it was extracted. Which system wins when two of them disagree. And is this use of it permitted here. None of those are properties of the data. They are properties of the data in use, which is why they only appear when a workflow is named.
The audit is not failing. It is answering correctly for the workload it was designed around. Reporting and analytics are batch, aggregate and retrospective. They tolerate overnight refreshes, they work on populations rather than individuals, and nothing breaks if a source is reconciled tomorrow rather than now.
A customer conversation is the opposite on all three counts. It is runtime, it concerns one individual, and it cannot wait for reconciliation. Data that is entirely fit for the first workload can be unusable in the second, and no measure of completeness or accuracy will reveal that, because completeness and accuracy are not the properties under strain.
Readiness is not a property of your data
If the quality scorecard is the wrong instrument, the obvious question is what the right one measures. It measures the data in use. Not whether a field is correct, but whether this workflow can reach it, trust it, resolve it to the right customer, settle it when systems disagree, and be permitted to use it at all.
That distinction has a practical consequence. The same field passes for one workflow and fails for another, on the same day, in the same estate. A balance that refreshes hourly is entirely fit for a spending summary and unfit for a fraud hold. An address that is correct in billing and stale in the CRM is harmless until an agent has to commit to a delivery date. Readiness cannot be certified once and applied everywhere, because the property being tested does not belong to the data. It belongs to the decision the data is being asked to support.
Data readiness for AI in customer experience is not an enterprise condition. It is a per workflow condition.
This is why the assessment has to come after the map. McKinsey's 2025 research found that organisations reporting significant financial returns from AI were about twice as likely to have redesigned their end-to-end workflows before selecting a modelling approach. Until you know which workflow the AI is performing, which decision sits inside it and which action follows, there is nothing specific enough to test. Workflow cartography comes first because data readiness has no unit of analysis without it.
What follows is the instrument itself. Five tests, applied per workflow, each returning a pass, a conditional pass or a fail.
The five tests of AI data readiness
Data readiness for a workflow resolves into five tests. Each returns a pass, a conditional pass or a fail. Not every test binds on every workflow, and testing tells you which of the five apply.
|
01
Reachability
Can the AI retrieve it, in the workflow, at the moment it needs it? Not does it exist. Can a system call it and get an answer back before the customer disengages.
|
02
Recency
Is it current enough for this decision? Yesterday's address is fine. Yesterday's appointment availability is not. The workflow sets the time budget, not the source.
|
|
03
Identity
Is this the same customer across every channel the workflow touches? Weak resolution either repeats questions the customer has answered, or joins two customers together.
|
04
Authority
When two sources disagree, which one wins, and is that rule written where a machine can read it? Mature estates fail this test most, because more systems means more places for one fact to live.
|
|
05
Permission
Is the AI authorised to use this data, for this purpose, in this channel, for this customer? Technically accessible is not the same as legally or ethically permissible. The consent flag usually lives in the marketing platform while the workflow runs in the service platform, so the agent cannot see the flag it is required to honour.
|
|
A workflow is only as data-ready as its weakest required test.
The word required matters. A post-authentication account query carries no identity risk. A workflow reading a single system of record has no authority conflict to resolve. Readiness is set by the worst of the tests that apply, not by an average across all five. Averages hide the failure that stops the workflow.
Each test is set out in full in the paper, with what fails it, how it is evidenced, and a worked telco example scored against all five. Download the full paper.
One customer, one question, four systems
A telecoms operator whose customer record is 99.4 percent complete, and which passes every governance test in the estate. An AI agent is asked, on a Tuesday afternoon, whether the customer's broadband will be live on Friday.
To answer, the agent needs the order status from the CRM, the provisioning state from the operational support system, the engineer appointment from field scheduling, and any open exception on the line from the port tracker. Four systems, one apparently simple question, asked thousands of times a week.
No precedence rule against the operational support system.
Overnight refresh only, and it disagrees with CRM order state.
Hourly refresh against a two hour promise.
No callable interface, screen only. Third party data, purpose scope unconfirmed.
|
Data quality
99.4%
Complete, governed, signed off by the data function. |
AI data readiness
Fail
Two failed tests and two unresolved authority questions. |
The same workflow, measured two ways, on the same day. Four sources, five tests, twenty answers. Fourteen are clean. The workflow still cannot run, because readiness is set by the six that are not. One of these numbers went to the steering committee. The other one decided whether the deployment worked.
Note what the verdict does not say. It does not say the workflow cannot be automated. It says three specific things have to happen first, and each of them is a piece of work with an owner and a duration. That is a scope. Without the assessment, it is an incident.
Fail does not mean stop
A failed test is not a two year data transformation. It is a named piece of work, and naming it is most of the value. A workflow that fails one test is a workflow with a date, once somebody owns the remediation.
The telco verdict on the previous section produces three items. Not a programme. Three items, each with an owner and a duration.
|
WEEKS
Port exception
Expose the port tracker through a callable interface, or route the exception check to a human step until it exists.
Integration, with network operations
|
DAYS TO WEEKS
Provisioning status
Move the refresh inside the promise window, or design the conversation to report status without committing to a date.
OSS owner, with conversation design
|
DAYS
CRM against OSS
Decide which system wins on order state, write the precedence rule with a divergence window, configure it.
Operations lead, with data
|
None of these is a data platform programme. Two of the three are decisions somebody already makes informally, written down for the first time.
That is why the cheapest remediation available is writing precedence rules. Every authority conflict you find is a rule that already exists in somebody's head. An advisor knows which system to believe. Getting that into a document, and then into configuration, costs days. It almost never appears on a data programme plan, because it is not a technology task and no technology function owns it.
Precedence is set per field, not per system. Billing can be authoritative for address and wrong about contact preference, at the same time.
Every rule needs a divergence window. Two sources that reconcile within an hour need a different rule from two that can disagree for two days.
Where no rule can be agreed, the workflow does not consume that field. Saying so explicitly is a design decision. Leaving it unresolved is a defect waiting for a customer to find.
A payment dispute workflow I worked on at a retail bank would have passed reachability, recency, identity and permission, and failed on authority alone. Dispute status existed in both the case management system and the card scheme portal, and the two could disagree for more than a day with no precedence rule between them. The remediation is a rule and a configuration change. Weeks, not quarters. Left unfound, a workflow like that either goes live contradicting the customer's own app, or sits behind a platform initiative that nobody sequenced against it.
Three patterns that recur
Three patterns recur across our assessments, consistently enough to plan around.
The failure is rarely accuracy
It is reachability or authority. The data exists, it is correct, and it either cannot be retrieved inside the time the conversation allows, or it exists in two systems that disagree with no rule for which one wins.
The interaction estate is the least ready and the most assumed
Organisations that would never deploy against unvalidated CRM data will point an AI system at ten years of unlabelled call transcripts and expect insight from it.
Readiness is the dimension most likely to be deferred
Remediation looks like infrastructure work, and infrastructure work is slow. Deferring it does not remove it. It relocates it, from a pre build workstream with a plan to a post go live incident with an audience.
| 95% |
of enterprise generative AI pilots delivered no measurable P&L impact
MIT Project NANDA, 2025
|
| 80%+ |
AI project failure rate, roughly twice that of non-AI IT projects
RAND Corporation, 2024
|
| 42% |
of companies abandoned most of their AI initiatives in 2025, up from 17% the year before
S&P Global Market Intelligence, 2025
|
| 2x |
more likely to have redesigned workflows before selecting a modelling approach, among organisations reporting significant returns
McKinsey, 2025
|
None of those studies measured model capability. They measured what happened around the model.
Value at stake × readiness = roadmap
One workflow is a decision. Forty workflows is a roadmap, and it is built from two numbers, not one. The five tests aggregate, and read across a portfolio they answer the question an executive actually has.
|
NOT READY · HIGH VALUE
Order status enquiry
Fails reachability. Needs a callable interface for port exceptions. Sequence it behind that work, do not abandon it.
|
NOT READY · HIGH VALUE
Payment dispute
Fails authority. Needs a precedence rule and a configuration change. Days of work standing in front of a high value workflow.
|
|
READY WITH CONSTRAINT · MEDIUM
Appointment rebooking
Recency is conditional. Design the conversation to avoid a two hour commitment and it can proceed now.
|
READY · MEDIUM VALUE
Billing explanation
No test binds. This is where a first wave starts, not because it is the most valuable, but because it is the one that can run.
|
Not which use cases are valuable, and not which are ready, but which valuable ones are ready, and precisely what stands between the rest and deployment. That overlay is where the roadmap comes from, and it is the subject of Paper 04.
Start with three workflows
Any capable operator can run a version of this. Take your three highest volume automation candidates and run the five tests against each, one workflow at a time. It takes under a week per workflow with the right people in the room, and the right people are not the data team alone. One person who owns the workflow operationally, one who knows the systems it touches, one who can speak to consent and retention. Write down how current the data has to be before anyone checks how current it actually is, because doing it in that order stops the available refresh rate from setting the standard.
Where an organisation wants an independent verdict, this is the discipline CXaiS applies. We test the workflows, name what fails and what it will take to fix, and hand back a readiness verdict alongside the value at stake. The verdict is honest because it has to be. A readiness assessment that clears a workflow it should not is worse than no assessment, because it puts a confident answer in front of a customer on data that could not support it.
The interaction estate
Everything so far concerns structured data. The larger and less examined half of the CX estate is the interaction record itself. Calls, chats, emails, messages, advisor notes, survey verbatims.
Organisations assume this estate is an asset because it is large. Volume is not readiness. Three problems recur.
There is a timing argument too. IBM's Institute for Business Value puts the average useful lifecycle of an AI model at approximately 14 months. An interaction estate that takes two years to label and structure is being prepared for a model generation that will be retired before the work completes. That argues for narrow, workflow-specific labelling ahead of general estate remediation, which is where the five tests point as well.
Data Readiness, in full
This page is the argument in summary. The paper runs it in full. Each of the five tests with what fails it and how it is evidenced, the telco workflow scored source by source, the remediation scope that comes out of a failed verdict, the interaction estate in detail, and the five maturity levels for this dimension, from Explore to Leading.
Twenty one pages, with the exhibits at full size.
Download the paper (PDF) ↓About and sources
Figures cited are drawn from published research as attributed. Observations described as ours are drawn from CXaiS client engagements and from the author's prior practitioner experience, and are stated as experience rather than research findings. This paper is published for information and does not constitute advice on any specific programme.
MIT Project NANDA, The GenAI Divide, 2025. RAND Corporation, The Root Causes of Failure for Artificial Intelligence Projects, 2024. S&P Global Market Intelligence, Voice of the Enterprise, AI, 2025. McKinsey, The State of AI, 2025. IBM Institute for Business Value, 2026.
|
See the hidden process before you automate it.
Start with an independent map of one workflow, upstream of any platform decision.
|
cxais.ai/contact |
