← All insights

Customer data maximization · 2

Why does identity resolution fail, and how do you fix it?

Identity resolution is the layer every other data project stands on, and the one most teams skip. Below 60% coverage your personalisation and attribution are guesses. Build the spine deterministically.

Why does identity resolution fail, and how do you fix it?

Identity resolution fails when teams treat it as a tool to buy rather than a decision to make. Most patch it with probabilistic matching and hope. You fix it by building the spine deterministically: verify email at capture, link by company domain and address, then cross-link every signal to one profile.

Identity resolution is the layer every other data project stands on. It is also the one most teams skip, because it is invisible to anyone outside the data team. The symptoms are not invisible.

What identity resolution failure looks like

Duplicate profiles. A customer who gets the welcome offer three times. Personalisation that misfires and reads as careless, because the system is talking to a fragment of a person, not the person.

None of this looks like an identity problem from the outside. It looks like a marketing mistake, or a sloppy campaign, or a bad vendor. Underneath, it is always the same thing. The business cannot reliably tell that two records are the same human being.

So it guesses. And below a certain coverage line, the guessing swamps the signal.

The evidence: below 60%, you are guessing

Below 60% identity coverage, your personalisation and your attribution are guesses. Not estimates. Guesses (Improvado, 2026).

That line matters because everything downstream inherits it. If you can only confidently link six in ten customers to themselves, then four in ten of your personalisation decisions are aimed at a ghost. Your attribution splits revenue across profiles that are secretly the same person. You optimise spend against numbers that are wrong in ways you cannot see.

Most teams respond by reaching for a probabilistic match or a vendor identity graph to patch the gap. That degrades as signal loss gets worse, and signal loss is getting worse every year as cookies die and consent tightens. Probabilistic matching is a supplement. It is not a spine.

Is this you?

Five checks. Yes or no.

  • Do some customers receive the same onboarding email more than once?
  • When you count “unique customers”, are you counting profiles instead of people?
  • Do you rely on a third-party graph you cannot audit to stitch identities together?
  • Can a customer log in on their phone and their laptop and be recognised as one person?
  • If a couple shares an account, do you know which of them actually did what?

If you are unsure on two or more, your identity spine is weaker than your dashboards suggest.

What weak identity costs

The cost is quiet, which is why it survives. It shows up as personalisation that annoys rather than helps, and as customers who feel the brand does not know them despite years of data.

It shows up as attribution you cannot trust, so budget flows to the channel that looks good on a broken measurement, not the one that works. It shows up as wasted sends against duplicates you are paying to store and mail. And it caps every AI project you try to run on top, because a model fed fragmented identities learns fragmented behaviour.

You do not see a bill. You see a business that is slightly worse at everything customer-facing than it should be.

How to build the spine deterministically

This is how I do it, and none of it needs a new platform.

Verify the email at capture. Real verification, double opt-in. Email is your most important key, so protect its quality at the front door, not after. A bad email poisons everything you later link to it.

Use company-domain email. Two people with the same company domain work at the same business. That is safe to infer, and useful for B2B and for household-versus-work separation.

Use precise address matching. A shared, precisely matched address usually means a household. That gives you the household view without guesswork.

Handle families with consent. Join members into one family account, but have one member invite the other. Consent stays clean, and you still get the household picture. Run this on your own B2C efforts too, not just the client’s.

Cross-link what you already hold. Email, cookie, device, in-app login, family, address. One person, one profile, many signals feeding it. Deterministic is the backbone. Probabilistic extends the edges, never carries the weight.

Capture more than one touchpoint, every time

Identity gets stronger the more keys you hold on the same person. So capture more than one, consistently: email, phone, SMS, address. Ask for both the company email and the personal one, and keep them apart. Every extra verified touchpoint is another way to recognise the same human when one key goes stale.

The personal email is the one to protect. Here is why. A high-value customer gives you their personal email, then changes jobs. Their work email dies. Their personal email does not. You still have them, and that is golden. Someone senior enough to be high-value rarely drops down. They move sideways or up, into a similar role with a bigger budget. The person who was worth reaching last year is worth more this year, and you are the one who can still reach them.

Company email tells you where someone works today. Personal email tells you who they are across every job they will ever hold. Capture both. Lean on the personal one for the long game.

Resolve identity on purpose

The teams that get this right do not have a better vendor. They made a decision. They chose to resolve identity deterministically, on purpose, instead of hoping a tool would do it quietly in the background.

That is the whole move. Decide that identity is a thing you build and own, verify your keys at the door, and cross-link with intent. Coverage climbs, the guessing shrinks, and every layer above it gets sharper.

Identity sits on top of a named master system. If your data is still scattered with no agreed source of truth, start one layer down with fragmented data and silos. And the keys you resolve on are only as good as their freshness, which is where data quality comes in.

Want the full operating sequence? Get the identity-resolution playbook: the three-step spine, the coverage metric to track, and how to train the team to hold it.

Go deeper on customer data maximization

Three ways forward. Pick the one that fits where you are.

  • Get the playbook. Practical notes on turning the customer data you already own into revenue, straight to your inbox. Join the newsletter at the foot of this page.
  • Take the assessment. Score your customer data maximization in four minutes and see your top revenue blockers. Start the assessment →
  • Book a meeting. Bring your data problem. Leave with a prioritised fix, not a platform pitch. Book a call →

Post 2 of 25 in the Customer Data Maximization series. Previous: Fragmented customer data. Next: Data quality, the AI multiplier.

Frequently asked questions

What is identity resolution?

Identity resolution is the process of linking every signal from one person to a single profile. Email, cookie, device, in-app login and address all point back to one customer. It is the layer that lets personalisation, attribution and a single customer view work at all.

What is a good identity resolution coverage rate?

Aim for identity coverage above 60% of your active customers. Below that line, personalisation and attribution stop being estimates and become guesses (Improvado, 2026). The higher your deterministic coverage, the more of your data can actually be acted on.

Deterministic or probabilistic identity resolution: which is better?

Deterministic matching is the spine because it uses keys you can trust: verified email, company-domain email, precise address. Probabilistic matching is a supplement that degrades as signal loss worsens. Use probabilistic to extend coverage, never as the backbone.

How do you resolve identity for families and households?

Join family members into one account, but have one member invite the other so consent stays clean. Precise address matching gives you the household view. This keeps personal and shared data separate while still letting you see the whole household.

Why does verifying email at capture matter so much?

Email is your most important key, so its quality decides the quality of everything linked to it. Verify at the point of capture with double opt-in, not in a cleanup six months later. Protecting the key at the front door is far cheaper than repairing it downstream.