Demo · runs entirely in your browser
Four fake systems (CRM, e-commerce, support desk, in-store) full of nicknames, typos, format drift, and missing fields — resolved into unified customers, grouped into mailable households, and enriched with a 3rd-party demographics file — right here in your browser, every merge explained, accuracy measured against known ground truth.
Field completeness across the 769 raw records:
| Customer | Phone | City |
|---|
This is the matching algorithm itself, run on real pairs from the dataset above — raw values, what normalization did to them, each field's similarity times its weight, and the verdict against the current threshold. Pick a case:
| Field | Weight | Similarity rule |
|---|---|---|
| 0.30 | exact match after lowercasing = 100%; otherwise Jaro-Winkler on the username, discounted (×0.75 same domain, ×0.4 different) | |
| Phone | 0.25 | exact match on normalized 10 digits, or nothing — phones don't do "close" |
| Name | 0.30 | 0.45 × JW(first) + 0.55 × JW(last), after nickname canonicalization (Bob → Robert); a bare initial can only score 80% on a first-letter match |
| Address | 0.15 | 0.6 × token overlap on the standardized street + 0.4 × ZIP exact |
score = Σ (similarity × weight) ÷ Σ weights, over the fields both records have — so a missing phone doesn't punish the pair, it just shifts the evidence burden. Then the family guards: clearly different full first names cap a pair at 70%, and a bare initial without email proof caps at 74%, both under the default threshold. Full source: match.js on GitHub — ~80 lines for everything on this page.
1. Normalize. Phones collapse to ten digits, emails validate, nicknames map to canonical names (Bob → Robert), street abbreviations standardize (Boulevard → Blvd), and malformed values are quarantined rather than silently matched.
2. Block. Comparing every record with every other would be ~295,000 pairs; cheap blocking keys (same phone, same email, ZIP + name prefix, a typo-tolerant name key) cut that to a few thousand candidate pairs without losing the true matches.
3. Score and explain. Each candidate pair gets a weighted similarity — exact email and phone matches, Jaro-Winkler on names (nickname-aware), token overlap on addresses — with weights renormalized over the fields both records actually have. Every merge above shows its per-field reasoning.
3b. Guard against families. Households are the classic trap: two siblings share a surname, an address, even the home phone. Two rules keep them apart — clearly different full first names cap a pair's score, and a bare initial cannot force a merge without email evidence. Genuinely ambiguous records stay unmerged on purpose; in production those route to a data-steward queue instead of a guess.
4. Cluster and survive. Pairs above the threshold link into clusters; each cluster becomes one golden record, with survivorship picking the most-attested, best-formed value for every field.
5. Household. Resolved customers at the same normalized address join a household when they share a surname or a phone — so the Smiths get one catalog, while unrelated roommates stay separate. Household grouping is scored against ground truth too.
6. Enrich. A synthetic 3rd-party demographics file appends onto golden records by normalized email or phone; corrupted vendor keys fail to match and are quarantined, not force-joined — consuming 1st- and 3rd-party data into one repository.
7. Measure honestly. Because the data is generated, the true customer identities are known — so precision and recall up top are real numbers, not vibes, and they move when you move the threshold. That tradeoff is the whole game in identity resolution and householding.
Honestly: this is a teaching demo, not a product — classic deterministic + probabilistic matching on a small synthetic dataset, built to show the concepts (and my code) end to end. Production identity resolution handles orders of magnitude more scale, streaming updates, and far messier reality.