← All insights

Customer data maximization · 22 · Playbook

AI data readiness: the full playbook

The operating sequence for making data ready before you scale AI. Audit readiness first, fix the foundation in order, then scale the model on data you trust.

AI data readiness: the full playbook

Making data ready for AI is a three-step sequence, run in order. Audit data readiness before you scope the model. Fix the foundation, identity then quality then unification, because there is no AI shortcut around them. Then scale the model on data you trust. Readiness is the precondition for AI spend, not a parallel workstream you bolt on later.

Most AI programmes run these steps backwards. They buy the model, then discover the data. This playbook runs them forwards, so the pilot ships instead of stalling.

Step 1. Audit readiness before you scope the model

Before a licence is bought or a data-science team is hired, look at the data the model will actually touch. Not a sample. The real thing.

Score three things, plainly. Is a customer resolved to one identity, or scattered across four systems under three spellings? Is the data fresh enough to act on, or decayed? Is there one agreed master, or several arguing copies? Write the answers down. That page is your readiness baseline.

This matters because data readiness for AI is the defining challenge of the moment, and most organisations are still catching up (IBM Institute for Business Value, 2025, surveying 1,700 leaders across 27 countries). The audit is what tells you whether you are one of them before you spend, not after.

Owner: the data leader, not the AI vendor. A vendor scoping your readiness will always find you ready enough to buy. What to measure: percentage of customers resolved to a single identity, data freshness by source, and whether a master system is named. Time: days, not months. That is the cheapest de-risking you will ever do.

Step 2. Fix the foundation, in order

AI does not transcend the data underneath it. Fed fragmented, unresolved, dirty data, AI perpetuates the inconsistencies at scale and worsens the problem it was meant to solve (Uniform, 2025). So fix the foundation first, and fix it in the order the layers depend on each other.

Unify the sources. Name one master system and connect everything to it, so a customer is one record rather than four. This is the base layer. Nothing above it works while the same person exists as three unlinked copies. The full sequence is in fragmented customer data, and how to fix silos.

Resolve identity. Link a person to themselves across email, device, and household using deterministic keys you can trust, not probabilistic guesses. A model cannot personalise for a customer it cannot recognise. The method is in identity resolution, and how to fix it.

Hold quality as a habit. Clean on collection, trade value for fresh data, and re-verify every six months. A model trained on dirty data scales your errors at machine speed, so quality is the multiplier on every AI project. The habit is in data quality, the AI multiplier.

Owner: a named person or team who owns the foundation across every touchpoint. What to measure: identity resolution rate rising, duplicate rate falling, data freshness holding between checks. There is no shortcut here, and any vendor who promises one is selling the stall you are trying to avoid.

Step 3. Scale the model on data you trust

Only now do you push from pilot to production. The order matters. A model scaled before the foundation is fixed scales the mess. A model scaled after it scales the result.

Re-run the pilot on real production data, the messy real thing you audited in Step 1, not the tidy sample. If it holds, the foundation is ready and you scale. If it breaks, the break tells you exactly which foundation layer still needs work, and you go back to Step 2 on that layer alone.

This is the discipline the firms scaling AI actually use. They fixed the data first, then let the model act on data they trust. The pilot that used to fall apart now ships, because the thing that broke it is gone.

Owner: the AI lead and the data leader together, signing off jointly that the foundation held on real data before scale. What to measure: pilot accuracy on production data matching accuracy on the sample, and the model in production making decisions the business can defend.

How to train your team to hold the fix

The sequence sticks only when the team stops treating the model as the project. That is a mindset change, and it needs teaching.

Start with the leadership framing. Say it plainly, in every AI kickoff: readiness is the precondition, not a parallel track. Make the readiness audit a gate the project cannot pass without. When a new AI idea arrives, the first question is not “which model” but “is the data ready”, and the audit answers it before money moves.

Then give the foundation an owner with standing. Identity, quality, and unification need a named person whose job survives the pilot. Without that, the foundation decays the moment the launch buzz fades, and the next model inherits the same mess.

Finally, reward the boring work. The team that keeps duplicate rates low and data fresh is doing the work that makes every future AI project cheaper and faster. Make that visible. The habit holds when the people who keep it are seen to keep it.

Where Morphy helps

We run a four-to-eight week AI data readiness sprint. It starts with the Step 1 audit: a plain scored view of whether your data is ready for the model you have in mind, and exactly which foundation layers are not.

From there we fix the highest-impact layer first, whether that is naming a master, raising the identity resolution rate, or standing up a quality habit. The engagement is tied to one defined metric agreed up front, usually the identity resolution rate on the customers your AI project will act on, or pilot accuracy holding on real production data.

You leave with a foundation the model can run on, not a slide deck about one. Every engagement ships a working improvement, not a recommendation to buy more tools.

Go deeper on customer data maximization

Three ways forward. Pick the one that fits where you are.

  • Get the playbook. Practical notes on turning the customer data you already own into revenue, straight to your inbox. Join the newsletter at the foot of this page.
  • Take the assessment. Score your customer data maximization in four minutes and see your top revenue blockers. Start the assessment →
  • Book a meeting. Bring your data problem. Leave with a prioritised fix, not a platform pitch. Book a call →

The playbook companion to Is your data ready for AI? Why most AI projects stall in pilot. Post 22 of 25 in the Customer Data Maximization series.

Frequently asked questions

What is the first step to AI data readiness?

Audit the data before you scope the model. Look at the data the model will actually run on and score three things: is a customer resolved to one identity, is the data fresh, and is there one agreed master. Treat readiness as the precondition for AI spend, not a workstream you run in parallel.

In what order should you fix a data foundation for AI?

Identity, quality, and unification are the work, and there is no AI shortcut around them. Unify sources so a customer is one record. Resolve the person to themselves across email, device and household. Keep records clean as a habit. Then, and only then, scale the model.

How do you know when data is ready for AI?

When a customer resolves to one identity, the data is fresh enough to act on, and one master system is agreed. Test it by running the model on real production data, not a hand-picked sample. If the pilot holds on the messy real thing, the foundation is ready.

Why does treating data readiness as a parallel workstream fail?

Because the model ships before the foundation is fixed, meets real dirty data, and stalls. Readiness has to lead, not run alongside. The firms scaling AI are the ones that fixed the data first, then scaled the model on top of it.