What does a text diff checker actually do?+
It works out the cheapest way to turn one text into the other. The only moves available are deleting a line and inserting a line, so a “difference” is really an entry in an edit script: this line went, this line arrived, everything else stayed. That framing explains most of what a diff does that surprises people. It has no notion of a line being edited (an edit is a deletion and an insertion that happen to sit next to each other), and it has no notion of a paragraph being moved, only of one disappearing here and appearing there. This page adds a second pass on top: when a deletion and an insertion are close enough to be versions of the same line, it compares them again inside and shows you the words that actually moved.
Why does the algorithm minimise edits instead of finding “the” differences?+
Because there is no such thing as “the” differences: there are only stories about how one text became the other, and infinitely many of them are consistent with the evidence. You could always describe any change as “delete everything, insert everything”, which is correct and useless. The useful constraint is minimality: prefer the story with the fewest edits, because it keeps the most of the original in place and therefore reads closest to what a person actually did. Myers' 1986 algorithm finds such a story in time proportional to the size of the input multiplied by the length of that script, which is why two nearly identical files compare in milliseconds however long they are.
Is the shortest edit script unique?+
No, and that is worth knowing before you argue with a diff. Compare “a b c” with “a c b”: you can delete the b and insert it after the c, or delete the c and insert it before the b. Both are two edits; neither is more correct. Every diff tool breaks the tie by convention, not by insight. This one biases towards reporting deletions before insertions and towards taking matches as early as possible. So when a diff attributes a change to a line you did not touch, it is usually not wrong, it has simply chosen a different equally-short account of the same change. Git's `--patience` and `--histogram` algorithms exist precisely to make that tie-breaking feel more human on source code.
Why does a block I moved show up as a deletion and an insertion?+
Because “move” is not one of the operations. The edit script has exactly two verbs, delete and insert, so relocating a paragraph is spelled as its removal from the old position and its arrival at the new one. That is twice the red and green for a change that felt like one action. Some tools bolt on move detection afterwards by looking for a deleted block that matches an inserted block elsewhere, but it is a heuristic layered on top, not part of the algorithm, and it goes wrong when the moved block was also edited. The practical workaround is to make moves and edits separate revisions: move the section in one pass, change its wording in the next, and each diff stays readable.
Why do whitespace-only changes flood the diff, and when is ignoring them a mistake?+
Because a line is compared as a whole string. Re-indenting a file, converting tabs to spaces, or letting an editor strip trailing spaces on save changes every line it touches, so a formatting pass and a genuine edit are indistinguishable until you tell the comparison to look past spacing. Turning whitespace off is the right move for reviewing prose or reformatted code. It is the wrong move for anything where spacing is syntax: Python's indentation defines block structure, YAML's indentation defines nesting, and a Makefile recipe must begin with a literal tab rather than eight spaces. In those files a whitespace-only change is not cosmetic. It is the change, and hiding it will hide a real bug. This page keeps the two ideas apart: “ignore leading and trailing whitespace” forgives the edges while still reporting indentation, and “ignore all whitespace” forgives everything.
Why do two files that look identical show every line as different?+
Almost always line endings. Windows ends a line with a carriage return followed by a line feed, Unix and macOS use the line feed alone, and older Mac software used the carriage return alone. The characters are invisible, so a file that travelled through a Windows editor can differ from its original on every single line while looking exactly the same on screen. This page splits on all three terminators and compares the line content, so CRLF against LF does not produce a false wall of changes. It does still tell you the two sides use different endings, because that difference is real and will matter to git, to a shell script's shebang, and to anything reading the file byte by byte. The other invisible culprits are a non-breaking space pasted from a web page and a byte-order mark at the very start of a file.
What does the @@ line in the unified patch mean?+
It is a hunk header, and it tells a patch program where the following block belongs. `@@ -12,7 +12,9 @@` reads: starting at line 12 of the original file, seven lines are described here; starting at line 12 of the new file, nine lines are described here. The minus range always refers to the original and the plus range to the revision, and the counts include the unchanged context lines shown either side, not just the changes. A count of one is written without its comma, and an empty range is written against the line it follows, so a pure insertion after line 4 appears as `-4,0`. Below the header, a leading space means unchanged, a minus means removed, and a plus means added. Context is what lets a patch apply to a file that has drifted slightly: with three lines either side, a patch program can still find the right place after unrelated edits above it.
Why does word-level highlighting need a similarity threshold?+
Because pairing a deleted line with an inserted line is a guess, and a bad guess is worse than no guess. If two lines really are versions of each other, comparing them inside is exactly what you want. If they merely happen to be adjacent, the inner comparison finds the handful of short words and stray punctuation they share and produces confetti: a scattering of green and red fragments that is harder to read than “this line went, this line arrived”. The threshold is a ratio: the characters the two lines share, counted on both sides, divided by their combined length. One means identical, zero means nothing in common. Above the setting you choose, the pair is refined and shown word by word; below it, the two lines are reported as a straight replacement. Loosen it when comparing heavily rewritten prose, tighten it when a diff starts looking like tinsel.
How big can the two texts be, and what happens at the limit?+
Each side takes up to 200,000 characters or 10,000 lines, which is roughly a 40,000-word manuscript. There are two softer limits inside that. The line comparison stops searching once the edit script would run past 12,000 edits or the search exceeds its effort budget, because comparing two texts with nothing in common is quadratic work and would otherwise lock the tab; past that point the affected region is reported as one block replaced by another, and a note says so. Separately, a single line longer than 4,000 words or characters is refined by shared opening and ending rather than fully, which is again flagged rather than hidden. Every one of these limits produces a correct answer, just a coarser one than usual.
Can this tell me whether two pieces of code do the same thing?+
No, and no diff tool can. This compares text, not meaning. Renaming a variable consistently, reordering two independent function definitions, swapping single quotes for double quotes, or reformatting an argument list across three lines are all large diffs and zero behavioural change. Conversely, `>` becoming `>=` is one character and can be the whole bug. A diff is evidence about what was typed, not about what will happen, which is exactly why code review reads diffs and then runs the tests. If you want a comparison that understands structure rather than characters, you need a tool that parses the language, and the useful first step here is to normalise the formatting of both sides before comparing so that the diff is about content rather than layout.