Copyable intake rule
Decide before a record becomes a duplicate.
| Evidence at intake | Route | Action |
|---|---|---|
| Same verified email or account ID | Exact match | Update the existing record and retain its owner |
| Similar company or domain, but a key field differs | Review | Show the likely match and the reason; do not merge automatically |
| Shared address, blank company, or conflicting company identifier | Block automatic matching | Create a labelled review case or request better evidence |
Start upstream
A duplicate record is usually created before anyone sees it.
Most teams clean duplicates after they have already damaged reporting, ownership, and follow-up. Each duplicate can spawn extra tasks, split communication history, and make pipeline numbers less trustworthy.
The better move is to treat intake as a gate. Before a lead, company, account, or contact becomes official, the system should check the few fields that reliably identify it and decide whether the new submission is a match, a possible match, or new.
Where new CRM records are created: forms, imports, integrations, manual entry, and enrichment tools.
Which fields are reliable enough to match on: email, domain, company number, account ID, or normalized company name.
Who owns ambiguous matches instead of letting every user make a different judgement.
What happens after a match: merge, update, reject, or route for human review.
Minimum rule set
Three rules catch most duplicate trouble without making intake heavy.
You do not need a perfect master data program to make progress. You need enough structure that the same account cannot quietly enter the workflow in four different shapes.
Normalize before matching
Lowercase emails, clean domains, remove legal suffix noise, and trim whitespace before comparing records. Most apparent duplicates are the same account written three different ways, normalization catches them without any fuzzy logic at all.
Separate exact and fuzzy matches
Exact matches can be handled automatically. Fuzzy matches should be routed to a review queue with the reason shown plainly, "same domain, different contact name" is actionable; a numeric confidence score is not.
Protect the owner field
Ownership should not be overwritten just because a duplicate or import arrives with a different assigned person. Silent ownership changes are how accounts fall through, the original owner stops seeing it and the new one does not know they have it.
The matching, one level down
Good fuzzy matching combines measures and looks for a reason to say no.
"Separate exact and fuzzy" is the rule; this is what fuzzy has to actually do. A single similarity score, taken alone, both misses real duplicates and invents fake ones. A few habits fix most of it.
Combine complementary measures. Edit distance alone rejects a reordered name: "John Smith" and "Smith, John" look far apart character by character, so pair it with a token-set measure that ignores word order. Each measure covers the other's blind spot.
Look for a reason to say no, not only yes. Any shared value can "prove" sameness if you only look for agreement. Before auto-merging, check for a disqualifier: two records with the same company name but different company numbers are two companies, not one.
Guard the placeholder. A shared info@ address or a blank company is not a match signal; block on it and you merge every contact at the company into one. Skip the values that thousands of rows share.
Leave the ambiguous ones unproven. A partial overlap should route to the review lane, not silently unlock a merge. Proven pairs merge; unproven pairs wait for a human.
Implementation path
A calm duplicate-control build usually follows four steps.
The aim is to make record creation predictable enough that sales, service, and reporting stop fighting the same data problem every week.
01
Inventory every intake source
List every place a record can be created or imported. Do not skip manual entry, because it is often where the quiet exceptions live.
02
Choose the matching keys
Pick the smallest set of fields that can reliably identify the record. If a field is often missing or messy, it should not carry the whole match.
03
Create a review lane
Ambiguous records should not silently create duplicates. Put them in a small queue with the likely match and the reason for uncertainty.
04
Measure recurrence
Track which source creates duplicates and which rule catches them. That shows whether the fix is working or one source still needs attention.
Before the gate
An intake gate stops new duplicates. It does not tell you how much mess already got in.
The gate is the forward fix. Sizing the backlog already sitting in the CRM is a separate, faster pass worth doing first.
Size the duplicates already inside before building the gate that blocks new ones.
Run a CRM data audit in a day ->In motion
What the intake gate surfaces as it runs.
Want the intake map checked?
Bring one intake source, the fields it sends, and a few recent duplicate examples. That is enough to see whether the fix is rules, routing, or a small build.
Bring the page, report, or workflow as it is now.
We reply with the clearest next step, or an honest no.
