Smart Deduplicator
Find near-duplicate rows that aren't exact matches — like "Jon Smith" vs "John Smith" — using real similarity scoring, not just exact comparison.
85% similar or more counts as a likely duplicate
How to use Smart Deduplicator
- Start by providing the required input: Drop a CSV file.
- Set the inputs and options shown here: Compare using column:, Similarity threshold.
- Follow the format note: 85% similar or more counts as a likely duplicate
- Run the tool with Find Near-Duplicates.
- Review the result or preview on the page and confirm it matches your input before using or sharing it.
Catching "Jon Smith" and "John Smith" as the same person
Exact-match deduplication misses the far more common real-world problem: the same entity entered slightly differently by different people — a typo, a middle initial, inconsistent spacing. Scoring text similarity between rows, rather than requiring an exact match, catches these near-duplicates that would otherwise slip through.
Choosing a sensible similarity threshold
A higher threshold (95%+) catches only very close near-duplicates with minimal false positives; a lower threshold (70-80%) catches more variation but risks grouping genuinely different entries together — start high and lower it gradually while reviewing results.
How is similarity actually calculated?
Using Levenshtein edit distance — essentially, how many single-character edits it takes to turn one value into the other, expressed as a percentage of similarity.
Which column should I compare on?
Whichever column best identifies a unique entity — typically a name, email, or company field rather than a numeric or date column.
Just need exact-match deduplication?
Use the Duplicate Row Remover instead — it's faster and simpler when you don't need fuzzy matching.