Poor data quality undermines AI because a model trained on bad data scales your errors instead of your revenue. Just 20% dirty data can cut model accuracy by around 10%. You fix it by cleaning on collection, trading value for fresh data, and re-verifying every six months, not by a one-off cleanup.
For years, poor data quality was a background tax. Annoying, expensive, tolerated. Then AI arrived and turned the tax into a multiplier.
Why AI changes the stakes on data quality
Poor data quality costs the average company 12.9 million dollars a year (Gartner). That was true before AI. It was survivable because humans in the loop caught the worst of it. Someone noticed the obviously wrong number and quietly fixed it.
AI removes that human pause. Feed a model bad data and it does not hesitate. It learns the errors, generalises them, and applies them to every decision at machine speed. Just 20% dirty data can cut a model’s accuracy by around 10% (FullStack, 2026). You do not get a slightly worse model. You get a confident one that is wrong more often, at scale.
That is the shift. Data quality used to cap how good your reporting was. Now it caps how good your AI is, and AI touches far more decisions than any dashboard ever did.
The cleanup trap
So most teams launch a big cleanup project. They scope it, resource it, run it for a quarter, and finish. Then they watch the data drift back to noise within months, because nothing changed about how data enters or ages.
This is the trap. A cleanup treats quality as a state you reach. It is not. Data decays from the moment you capture it, because people change jobs, move house, and abandon addresses. A record you perfected in January is drifting by March. Quality is not a place you arrive. It is a habit you keep.
Is this you?
Five checks. Yes or no.
- Have you run a big data cleanup, only to watch the data degrade again?
- Do you clean data in batches, months after it was captured?
- Are you holding off on an AI project until the data is “ready”?
- Do you collect email and phone but give customers no reason to keep them current?
- When did you last confirm a customer still wants to hear from you?
If you recognise three of these, your data is decaying faster than you are maintaining it.
Four moves that hold quality over time
This is what I tell clients. None of it is a heroic project. All of it is a habit.
Start anyway. Do not let data quality stop you from starting. Use what you can, while it is fresh. Data has a half-life, and a clean record today is worth more than a perfect record next quarter. Waiting for perfect data is how good years get wasted.
Clean on collection. The cheapest place to fix data is the moment it enters. Verify at capture. A wrong email caught at the door costs nothing. The same wrong email, discovered six months later after it has spread across four systems, costs a cleanup project.
Trade value for it. People share data when they get something back. Give a discount, a freebie, or points in return for the customer telling you more about themselves. Capture email, phone, SMS and messaging with permission, and give real value for the permission. Quality you paid for with value beats quality you scraped and hoped for.
Re-verify every six months. Stand up a preference centre and use it. Ask, twice a year, whether they are still there and still want to hear from you. This keeps consent clean and keeps the data fresh, and it does both without a cleanup sprint. Quality becomes a rhythm, not an emergency.
You do not need perfect data
You need fresh, permissioned data that you actually use. That is a different target, and a reachable one.
Perfect data is a fantasy that keeps teams from starting. Fresh, permissioned, in-use data is an operating discipline that compounds. The companies pulling ahead with AI are not the ones with the cleanest historical database. They are the ones with the fastest habit of capturing good data, using it while it is warm, and re-checking it before it rots.
Quality is the multiplier on every AI project you will run. Get the habit right and AI scales your revenue. Get it wrong and AI scales your mistakes. There is no third option.
This depends on the layers beneath it. Fresh keys are worthless if you cannot link them to a person, which is identity resolution, and neither works if your customer data is scattered across silos with no master.
Want the full operating sequence? Get the data-quality playbook: the three-step habit, the preference-centre re-verify loop, and how to train the team to hold it.
Go deeper on customer data maximization
Three ways forward. Pick the one that fits where you are.
- Get the playbook. Practical notes on turning the customer data you already own into revenue, straight to your inbox. Join the newsletter at the foot of this page.
- Take the assessment. Score your customer data maximization in four minutes and see your top revenue blockers. Start the assessment →
- Book a meeting. Bring your data problem. Leave with a prioritised fix, not a platform pitch. Book a call →
Post 3 of 25 in the Customer Data Maximization series. Previous: Identity resolution. Related: The half-life of data.
Frequently asked questions
How does poor data quality affect AI models?
Poor data quality multiplies errors. Just 20% dirty data can cut a model's accuracy by around 10% (FullStack, 2026). A model trained on bad data scales your mistakes at machine speed instead of scaling your revenue. Quality is the multiplier on every AI project.
How much does poor data quality cost?
Poor data quality costs the average company roughly 12.9 million dollars a year (Gartner). The cost lands as wasted spend, misfired campaigns, wrong decisions and degraded AI. AI raises the stakes because it acts on bad data faster and at larger scale than any human team.
Should you fix all your data before starting an AI project?
No. Do not let data quality stop you from starting. Use what you can while it is fresh, because data has a half-life. A clean record today is worth more than a perfect record next quarter. Start with the data that is good enough to act on now.
When is the cheapest time to clean data?
At the moment of collection. Cleaning on capture, with verification at the point of entry, is far cheaper than a batch cleanup six months later. The longer bad data sits and spreads across systems, the more it costs to fix and the more damage it does downstream.
How do you keep data quality high over time?
Treat it as a standing habit, not a project. Trade value for better data by giving customers a discount, a freebie or points in return for telling you more. Then re-verify every six months through a preference centre that confirms they are still there and still want to hear from you.