Toolkit
All tools
List cleanup · Free

Remove Duplicate Lines

Paste a list and get it back without the repeats, or with only the repeats. Case, stray spaces, punctuation, and accents are each a separate matching switch, and the lines that survive are written out exactly as you typed them.

Deduplicated in your browser · Nothing uploaded

5 lines out of 9, with 4 removed across 3 repeated lines.

Your list

Paste one item per line

9 lines in, 0 of them blank.210 / 200,000
Lines out
5
Removed
4
Repeated lines
3
Distinct
5
115 characters out4 duplicate copies found0 blank lines
What repeated

3 lines appear more than once

Sorted by how often each one occurred. Where two spellings compared as the same line, the other spellings are listed underneath, which is usually where a stray space or a capital letter turns up.

Each repeated line with the number of times it appeared
TimesLineFirst seen
3[email protected]line 1
2[email protected]Also matched: “[email protected] ”line 2
2[email protected]line 3
How it works

One pass, a hash map, and a rule about what not to touch.

Each line is turned into a comparison key by applying whichever matching rules you have chosen, and the keys go into a map that remembers the first line that produced each one, how many lines followed it, and which other spellings arrived under the same key. That is a single pass over the list rather than a comparison of every line against every other, which is why fifty thousand rows stay instant. The important part is what happens to the text itself: nothing. The key is used for matching and then discarded, and the output is your original line, byte for byte. Options that change your data and options that change your comparison are different things, and mixing them up is how a cleanup turns into a data loss.

  1. 01

    Paste the list, one item per line

    Emails, tags, file names, URLs, SKUs, or anything else that arrives one per row. Blank lines are dropped by default, and the character and line counts update as you paste so you can see immediately whether the whole export made it in.

  2. 02

    Decide when two lines count as the same

    Case, surrounding spaces, repeated inner spaces, punctuation, and accent normalisation are separate switches. An invisible trailing space is the single most common reason a duplicate survives a naive dedupe, so that one is on by default.

  3. 03

    Choose what survives, then sort it

    Keep the first copy, keep the last, keep only lines that appeared exactly once, or flip it round and keep only the lines that repeated. Then sort alphabetically, naturally so item2 comes before item10, by length, or by how often each line occurred.

For exports, mailing lists, tag sets, and file listings

Dedupe, invert it, sort it naturally, and see what repeated.

Matching options never rewrite your text

Ignoring case, spaces, or punctuation changes only what counts as a duplicate. Whatever survives is written out exactly as you typed it, with its original capitals and spacing intact. A deduplicator that hands back lowercased text has silently edited your data instead of filtering it, and that is usually discovered much later.

Four ways to keep, including the inverse

Keep the first occurrence, keep the last, keep only the lines that appeared exactly once, or keep only the lines that repeated. That last mode is the one most tools leave out, and it is what you want when the duplicates are the finding: repeated order ids, double-booked names, a term that appears twice in a glossary.

Natural sort, so item2 comes before item10

Plain alphabetical sorting puts report-10 before report-2, because it compares character by character. Natural sorting reads the digits as numbers, which is almost always what a list of file names or versions actually wants. Both are offered, along with sorting by length and by frequency.

The repeat table shows the near misses

Every line that occurred more than once is listed with its count and its first position. Where two different spellings compared as the same line, the other spellings are printed underneath, so a stray capital or a trailing comma is visible rather than merely absorbed.

Accent normalisation that catches invisible differences

Text copied from macOS filenames is frequently in Unicode NFD, where an accented character is stored as a letter plus a combining mark. It looks identical to the NFC form and compares unequal. Normalising to NFC before matching catches that class of duplicate, which is otherwise close to impossible to spot by eye.

Counts you can paste into the result

Turn on counts and each output line is prefixed with how many times it appeared, turning a deduplicated list into a frequency table you can paste straight into a spreadsheet. Numbering is a separate switch for when the list is going into a document.

List questions

Preserving order, invisible characters, emails, and finding repeats.

How do I remove duplicate lines without losing the order?+

Use the default, which keeps the first occurrence of each line and leaves the sort on original order. Every line stays where it first appeared and only the later copies are dropped. That matters more than it sounds: a list where order carries meaning, such as a playlist, a build sequence, or a set of steps, is ruined by a deduplicator that sorts as a side effect. If you would rather keep the last occurrence, for instance when later rows in an export are the more recent version of a record, switch to keeping the last.

Why did an obvious duplicate survive?+

Almost always an invisible character. A trailing space, a non-breaking space pasted from a web page, a stray comma left over from a CSV column, or a difference in capitalisation. Turn on “Ignore surrounding spaces” and turn off “Case matters” and most of them collapse immediately. If a pair still refuses to match, look at the repeat table: the entries listed under “also matched” show which spellings were treated as the same line, which usually makes the difference obvious. Genuinely different Unicode characters that look alike, such as a curly apostrophe against a straight one, will not match, and should not.

Can I find the duplicates rather than remove them?+

Yes, that is the “only the repeated lines” mode. It shows one row per line that occurred more than once, and the table underneath gives the count and the first position of each. It is the right mode for checking an export for double entries, finding a keyword that appears twice in a list that is supposed to be unique, or confirming that a merge did not introduce copies. The opposite mode, “only lines that appear once”, discards every repeated line entirely, including its first copy, which is what you want when a repeat means the row is suspect.

What is natural sorting?+

Sorting that reads runs of digits as numbers rather than as characters. Ordinary alphabetical sorting compares position by position, so “report-10” comes before “report-2” because the character 1 sorts before 2. Natural sorting compares 10 against 2 as numbers and puts them the right way round. It is what file managers use, and it is almost always what you want for anything with a numeric suffix: versions, chapters, invoice numbers, image sequences.

Does this work for a list of emails?+

Yes, and email addresses are the most common use for it. Two points worth knowing. Domains are case-insensitive by specification, so it is safe to leave “Case matters” off, and doing so catches the very common [email protected] against [email protected] pair. Local parts, the part before the at sign, are technically case-sensitive, but in practice essentially no provider treats them that way. Gmail's dots and plus-addressing are a different matter: [email protected], [email protected], and [email protected] all reach the same inbox but are genuinely different strings, and this tool will not merge them, because doing so would be wrong for every other provider.

How many lines can it handle?+

Up to 200,000 characters or 50,000 lines, whichever comes first, and it stays responsive because the work is a single pass over the list with a hash map rather than a comparison of every line against every other. A list past that ceiling is truncated and the page says so rather than silently processing part of it. For a very large export, split the file and run it in passes, or dedupe it where it lives with sort and uniq on a command line.

Are blank lines treated as duplicates of each other?+

If you leave them in, yes: every empty line is the same as every other empty line, so a run of them collapses to one. That is usually the desired behaviour for a list. If blank lines are meaningful in your text, for example because they separate paragraphs, this is the wrong tool for that content, since the whole model here is that a line is an item.

Does the list get uploaded?+

No. The comparison, the sorting, and the counting all happen in this tab, and nothing is transmitted or stored. That is the main reason to use a page like this rather than pasting a customer list, an internal URL set, or an export of user emails into a service that processes it on a server.

What is the difference between distinct lines and lines out?+

Distinct is how many different lines exist in your input under the current matching rules. Lines out is how many the current keep mode actually writes. They are the same when you keep the first or the last copy of everything, and they differ in the other two modes: keeping only lines that appear once excludes every repeated line, and keeping only repeats excludes everything that appeared once. Comparing the two numbers is a quick way to see how much of your list was duplicated.

Can I sort without removing anything?+

Not in this tool, and deliberately so. Deduplicating is what it does, so every line that survives is unique under your matching rules. If you want to sort a list with its repeats intact, turn on counts instead: you get one row per distinct line with the number of occurrences beside it, which carries the same information in a form that is easier to read and to paste into a spreadsheet.

More focused tools, ready when you are.

Explore the growing collection for calculations, documents, writing, and everyday work.

Browse all tools