How do I remove duplicate lines without losing the order?+
Use the default, which keeps the first occurrence of each line and leaves the sort on original order. Every line stays where it first appeared and only the later copies are dropped. That matters more than it sounds: a list where order carries meaning, such as a playlist, a build sequence, or a set of steps, is ruined by a deduplicator that sorts as a side effect. If you would rather keep the last occurrence, for instance when later rows in an export are the more recent version of a record, switch to keeping the last.
Why did an obvious duplicate survive?+
Almost always an invisible character. A trailing space, a non-breaking space pasted from a web page, a stray comma left over from a CSV column, or a difference in capitalisation. Turn on “Ignore surrounding spaces” and turn off “Case matters” and most of them collapse immediately. If a pair still refuses to match, look at the repeat table: the entries listed under “also matched” show which spellings were treated as the same line, which usually makes the difference obvious. Genuinely different Unicode characters that look alike, such as a curly apostrophe against a straight one, will not match, and should not.
Can I find the duplicates rather than remove them?+
Yes, that is the “only the repeated lines” mode. It shows one row per line that occurred more than once, and the table underneath gives the count and the first position of each. It is the right mode for checking an export for double entries, finding a keyword that appears twice in a list that is supposed to be unique, or confirming that a merge did not introduce copies. The opposite mode, “only lines that appear once”, discards every repeated line entirely, including its first copy, which is what you want when a repeat means the row is suspect.
What is natural sorting?+
Sorting that reads runs of digits as numbers rather than as characters. Ordinary alphabetical sorting compares position by position, so “report-10” comes before “report-2” because the character 1 sorts before 2. Natural sorting compares 10 against 2 as numbers and puts them the right way round. It is what file managers use, and it is almost always what you want for anything with a numeric suffix: versions, chapters, invoice numbers, image sequences.
Does this work for a list of emails?+
Yes, and email addresses are the most common use for it. Two points worth knowing. Domains are case-insensitive by specification, so it is safe to leave “Case matters” off, and doing so catches the very common [email protected] against [email protected] pair. Local parts, the part before the at sign, are technically case-sensitive, but in practice essentially no provider treats them that way. Gmail's dots and plus-addressing are a different matter: [email protected], [email protected], and [email protected] all reach the same inbox but are genuinely different strings, and this tool will not merge them, because doing so would be wrong for every other provider.
How many lines can it handle?+
Up to 200,000 characters or 50,000 lines, whichever comes first, and it stays responsive because the work is a single pass over the list with a hash map rather than a comparison of every line against every other. A list past that ceiling is truncated and the page says so rather than silently processing part of it. For a very large export, split the file and run it in passes, or dedupe it where it lives with sort and uniq on a command line.
Are blank lines treated as duplicates of each other?+
If you leave them in, yes: every empty line is the same as every other empty line, so a run of them collapses to one. That is usually the desired behaviour for a list. If blank lines are meaningful in your text, for example because they separate paragraphs, this is the wrong tool for that content, since the whole model here is that a line is an item.
Does the list get uploaded?+
No. The comparison, the sorting, and the counting all happen in this tab, and nothing is transmitted or stored. That is the main reason to use a page like this rather than pasting a customer list, an internal URL set, or an export of user emails into a service that processes it on a server.
What is the difference between distinct lines and lines out?+
Distinct is how many different lines exist in your input under the current matching rules. Lines out is how many the current keep mode actually writes. They are the same when you keep the first or the last copy of everything, and they differ in the other two modes: keeping only lines that appear once excludes every repeated line, and keeping only repeats excludes everything that appeared once. Comparing the two numbers is a quick way to see how much of your list was duplicated.
Can I sort without removing anything?+
Not in this tool, and deliberately so. Deduplicating is what it does, so every line that survives is unique under your matching rules. If you want to sort a list with its repeats intact, turn on counts instead: you get one row per distinct line with the number of occurrences beside it, which carries the same information in a form that is easier to read and to paste into a spreadsheet.