Toolkit
All tools
Pattern workbench · Free

Regex Tester & Debugger

Write a pattern, paste the text you are really working against, and watch every match light up as you type, with capture groups, named groups, offsets, a replacement preview, and syntax errors explained in English rather than in the runtime's own words.

Your browser's own RegExp · Nothing uploaded

Regular expression workspace

3 matches, 3 capture groups. First match at index 9.

Pattern

Write it without the slashes

Or paste a whole literal such as /\d+/gi and the flags come with it. Everything below recalculates on every keystroke.

g

Compiles. 3 capture groups (year, month, day) · pattern is 70 characters.

Flags
Matches in the test text
3

/\b(?<year>\d{4})-(?<month>0[1-9]|1[0-2])-(?<day>0[1-9]|[12]\d|3[01])\b/g against 114 characters.

Matches
3
Capture groups
3
Named groups
3
Zero-width hits
0
Text covered
26.3%
Run time
–
Test text

The subject the pattern runs against

Matches tint as you type. Adjacent matches alternate shade and each one carries a rule down its left edge, so two touching hits never read as one.

114 / 100,000 characters · 4 lines

Results

Every match, with its offsets and its capture groups.

Match #1 · capture groups
index 9 to 19
  • $1year2026
  • $2month07
  • $3day27

Turn on the d flag to see where each group sits, not just what it holds.

Preset library

Patterns that are honest about what they match

Each one loads its pattern, its flags, a replacement, and sample text containing the cases it deliberately rejects.

Shape checks

Anchored patterns that ask whether a whole line is the right shape.

Finding things in text

Unanchored patterns that pull items out of running prose or logs.

Rewriting

Patterns whose point is the substitution rather than the search.

Cheatsheet

Every token, what it means, and the smallest example that shows it

JavaScript syntax specifically: the notes call out the places it parts company with PCRE and with Python.

Character classes

One position in the subject, and the set of characters allowed to fill it.

.
Any character except a line break, unless the s flag is on./a.c/ matches “abc”, “a-c”, not “a\nc”
[abc]
Any one of the listed characters./[aeiou]/ on “rhythm” finds nothing
[^abc]
Any one character not listed. The caret must be first./[^0-9]/ on “a1” matches “a”
[a-z]
A range, by code point order./[a-f0-9]/ is one hex digit
\d \D
An ASCII digit, and anything that is not one. \d is exactly [0-9]; it does not include Arabic-Indic or fullwidth digits./\d{4}/ on “year 2026” matches “2026”
\w \W
A word character [A-Za-z0-9_], and its complement. Accented letters are not word characters./\w+/ on “café” matches “caf”
\s \S
Whitespace (space, tab, newline, form feed, NBSP, and the Unicode space separators), and its complement./\s+/ collapses runs of blanks
\p{…}
A Unicode property. Needs the u flag; without it \p means a literal “p”./\p{Lu}/u matches one uppercase letter in any script

Quantifiers

How many times the thing in front may repeat. Greedy by default; add ? to make it lazy.

*
Zero or more. Always succeeds, which is why /a*/g matches everywhere./ab*c/ matches “ac” and “abbbc”
+
One or more./\d+/ on “abc123” matches “123”
?
Zero or one: the atom is optional./colou?r/ matches both spellings
{3}
Exactly three times./\d{3}/ matches “555”
{2,}
Two or more times, no upper bound./o{2,}/ on “ooooh” matches “oooo”
{2,4}
Between two and four times./a{2,4}/ on “aaaaa” matches “aaaa”
*? +? ?? {n,m}?
The lazy forms: take as little as possible and grow only if the rest of the pattern fails./<.+?>/ on “<a><b>” matches “<a>”, not “<a><b>”

Anchors and boundaries

Zero-width assertions. They consume nothing and only ask about position.

^
Start of the subject, or start of any line with the m flag./^ERROR/m finds every line that begins with ERROR
$
End of the subject, or end of any line with the m flag./;$/m finds lines ending in a semicolon
\b
A word boundary: between a \w and a non-\w. Defined on ASCII word characters, so it lands mid-word in “naïve”./\bcat\b/ does not match “category”
\B
Not a word boundary./\Bcat/ matches the “cat” inside “bobcat”

Groups, alternation, references

Capturing is what turns a test into extracted data, and what costs memory when you did not need it.

(…)
Capture group, numbered left to right by opening bracket./(\d+)-(\d+)/ gives two groups
(?:…)
Group without capturing. Same grouping, no slot in the result./(?:ab)+/ repeats “ab” and captures nothing
(?<name>…)
Named capture group, ES2018. Available as match.groups.name./(?<year>\d{4})/ then match.groups.year
|
Alternation, tried left to right. The first branch that lets the whole pattern succeed wins./cat|category/ on “category” matches “cat”
\1
Backreference: the exact text group 1 already matched./(\w+) \1/ finds a doubled word
\k<name>
Backreference to a named group./(?<q>['"]).*?\k<q>/ matches balanced quotes

Lookaround

Assertions that inspect the text around the current position without consuming it.

(?=…)
Lookahead: what follows must match./\d+(?= kg)/ on “80 kg” matches “80”
(?!…)
Negative lookahead: what follows must not match./\bfoo(?!bar)/ skips “foobar”
(?<=…)
Lookbehind: what precedes must match. ES2018, and unlike most engines JavaScript allows it to be variable-length./(?<=£)\d+/ on “£40” matches “40”
(?<!…)
Negative lookbehind: what precedes must not match./(?<!un)happy/ skips “unhappy”

Flags

Set on the regex as a whole. JavaScript has no inline (?i) modifier: the flag applies everywhere or nowhere.

g
Global: every match, and lastIndex carries between exec calls./a/g with matchAll
i
Case-insensitive./hello/i matches “HELLO”
m
Multiline: ^ and $ match at line breaks. Does not affect the dot./^\s*$/m finds blank lines
s
Dot-all: . matches line breaks too./<p>.*<\/p>/s spans lines
u
Unicode mode: code points, \p{…}, \u{…}, and strict escapes./^.$/u matches a single emoji
y
Sticky: match only at lastIndex, do not scan forward.Tokenisers step lastIndex by hand
d
Indices: match.indices holds [start, end] for every group.match.indices[1] locates group 1

Escapes and literals

Twelve characters are special. To match one literally, put a backslash in front of it.

\. \* \+ \?
The metacharacters, taken literally. The full set is . * + ? ^ $ { } ( ) | [ ] \/\d+\.\d+/ matches “3.14”
\\
A literal backslash. In a string literal it doubles again to \\\\.new RegExp("\\\\d") is /\\d/
\n \r \t
Line feed, carriage return, tab./\r?\n/ splits either line ending
\xFF
A character by two-digit hex code./\x41/ matches “A”
\u00e9
A character by four-digit hex code./\u00e9/ matches “é”
\u{1F600}
A code point above U+FFFF. Requires the u flag./\u{1F600}/u matches one emoji
\0
The NUL character. Not an octal escape: those are illegal under u./\0/ matches U+0000

This is your browser's engine

The pattern goes straight to the runtime's RegExp constructor, so what you see is what Node and the browser will do. Python's re, PCRE, Go's RE2, and grep all differ: sometimes in syntax, sometimes in what identical syntax means.

\d is [0-9] here, and not everywhere

JavaScript keeps \d, \w, and \b on ASCII regardless of the u flag. Python and .NET widen \d to every Unicode decimal digit by default, so ٢٠٢٦ passes there and fails here. Ask for \p{Nd} when you want the wide meaning.

Nesting is out of reach

No regular expression can match balanced brackets, which rules out HTML, JSON, and source code however long the pattern grows. The HTML preset here is deliberately labelled as a skim, not a parser.

Everything on this page runs in your browser against your browser’s own RegExp implementation: no pattern, no test text, and no result is ever sent anywhere. The guards are stated rather than implied: up to 100,000 characters of test text, 5,000 matches, and 250 ms of wall clock per run, after which the run stops and says so. A zero-length match always advances the cursor by a full code point, so a pattern like /a*/g terminates instead of looping.

How it works

One pattern, one subject, and everything the engine knows about the pair.

A regular expression is a tiny program, and the reason it is hard to debug is that it runs invisibly: you see the answer, never the working. This page makes the working visible. Your pattern is compiled by the browser's own RegExp constructor, run over your text under a cap on length, match count, and wall-clock time, and reported match by match: where each one starts and ends, what every capture group holds, which groups did not participate at all, and what a substitution would produce. Because it is the runtime's engine and not a reimplementation, the results are the results your JavaScript will get. They are not, honestly, the results Python, PCRE, Go's RE2, or grep will get.

  1. 01

    Write the pattern and choose the flags

    Type the pattern without the surrounding slashes, or paste a whole literal such as /\d+/gi and the flags come across with it. Each of the seven flags is a labelled switch rather than a letter to remember, and a syntax error is answered in a sentence instead of the runtime's phrasing.

  2. 02

    Paste the text you are actually working against

    Real input, not a contrived example: a log excerpt, a CSV column, a page of prose. Matches highlight as you type, alternating shade so two matches that touch are still two matches, and a zero-width match is drawn as a marker rather than disappearing.

  3. 03

    Read the groups, then try the substitution

    Every match lists its offsets and each numbered and named group, including the ones that did not participate. Switch to Replace to preview a substitution with $1, $<name>, $& and friends expanded, or to Split to see what String.split makes of the same pattern.

Built for debugging, not for demonstrating

The parts other testers round off: empty matches, absent groups, and patterns that fight back.

Matching that cannot hang the tab

A zero-length match advances the cursor by a whole code point, so /a*/g terminates instead of looping forever. A wall-clock budget stops a run that has started grinding and says the pattern is backtracking badly, rather than freezing behind a spinner that never resolves.

Groups that did not participate are still shown

On /(a)|(b)/ against “b”, group 1 is undefined and group 2 holds the match. Testers that drop the empty slot silently renumber everything after it. Here the slot stays, labelled as not having participated, which is exactly the distinction that breaks code downstream.

Offsets for every group, not just the match

Turn on the d flag and each capture reports its own start and end index, read from match.indices. Without it you get the position of the whole match and have to search for the group's text again, which goes wrong the moment the same text appears twice.

A replacement preview that expands the awkward cases

$1, $<name>, $&, $` and $' and $$ are all expanded here rather than described, including the rule that $12 means group 12 when it exists and group 1 followed by a “2” when it does not. The output is checked against what String.replace produces.

Syntax errors translated into something actionable

“Invalid regular expression: /(/: Unterminated group” becomes “A ( was opened and never closed”, with a note that a literal bracket needs escaping. Around fifteen error families are recognised across the wording that V8, JavaScriptCore, and SpiderMonkey each use.

Ten presets that say what they really match

The email pattern is described as the pragmatic form that catches typos, not as RFC 5322 validation. The IPv4 one bounds each octet to 255 and admits it will still find an address inside a version string. Every preset ships with sample text containing the cases it rejects.

Regex questions

Greedy against lazy, catastrophic backtracking, and where JavaScript parts company with PCRE.

Which regex flavour does this test?+

The one already running in your browser. There is no engine bundled into this page: your pattern is handed to the JavaScript runtime's own RegExp constructor and the matches you see are the matches your code will get. That makes the results exactly right for JavaScript, TypeScript, Node, Deno, and Bun. For anything else they are only approximately right. Python's re, PCRE in PHP and Perl, Go's RE2, Rust's regex crate, .NET, Java, grep, sed, and ripgrep all differ, sometimes in syntax and sometimes in what the same syntax means. If the pattern is destined for a different language, test it there before shipping it.

What is the difference between greedy and lazy quantifiers?+

A greedy quantifier takes as much as it possibly can and then gives characters back one at a time until the rest of the pattern fits. A lazy one (written by adding a question mark, so *? +? ?? {n,m}?) takes as little as it can and grows only when forced. The classic demonstration is /<.+>/ against “<a><b>”: the greedy .+ swallows everything to the end, backtracks to the last “>”, and returns “<a><b>” as a single match. Change it to /<.+?>/ and you get “<a>”, then “<b>”. Neither is more correct; they answer different questions. The practical rule is that greedy is the right default when the delimiter is unique and lazy is the right default when it repeats.

What is catastrophic backtracking?+

It is what happens when a pattern can split the same text many different ways and the match ultimately fails. JavaScript matches by backtracking, so on failure it retries every remaining combination before giving up. The textbook trigger is a repeat inside a repeat, such as /(a+)+$/ against a long run of “a” followed by one “b”. There are exponentially many ways to divide 30 a's into groups of one or more, and the engine tries all of them before concluding the “$” cannot match. Thirty characters take milliseconds; forty take about a thousand times longer. The fixes are to make the inner and outer repeats unable to match the same text, to replace a quantified group with a character class, or to anchor the pattern so failure is detected early. This page scans your pattern for those shapes and names them before you run it.

Why did lookbehind arrive so late in JavaScript?+

Lookahead was in the language from the beginning, but lookbehind only landed in ES2018 (more than twenty years later), and until Safari shipped it in 2023 a pattern using it threw a syntax error on a browser many sites still supported. The delay was partly the general slowdown in language evolution between ES4's abandonment and ES6, and partly that lookbehind is genuinely harder to implement: the engine has to match backwards from the current position. JavaScript's implementation ended up better than most, because unlike PCRE, Python's re, and Java it allows variable-length lookbehind. /(?<=\$\d+ )item/ is legal here and rejected outright in those. If a pattern of yours needs to run in an old environment, the traditional workaround is to capture the preceding context in a group and discard it afterwards.

What is the difference between the u and v flags?+

The u flag, from ES2015, puts the pattern into Unicode mode: astral characters such as emoji count as one unit rather than two UTF-16 halves, \u{1F600} becomes valid, \p{…} property escapes are unlocked, and escapes with no meaning become syntax errors instead of being quietly ignored. The v flag, from ES2024, is a superset that adds set notation inside character classes (difference with --, intersection with &&, and multi-character string properties), so you can write [\p{Letter}--[a-z]] to mean “any letter except an ASCII lowercase one”. The two flags are mutually exclusive: a regex is either u-mode or v-mode, never both. This page offers u and not v, because a pair of buttons that silently cancel each other is a worse interface than one that is honest about its scope.

Why does my global regex skip every other match?+

Because a regex carrying the g or y flag is stateful. It stores a lastIndex property, and exec and test both start from there and update it on success. Store one in a module-level constant, call test twice on the same string, and the second call starts halfway through and returns false. The same bug bites when a global regex is reused across items in a loop. There are three ways out: build the regex fresh where it is used, reset lastIndex to 0 before each call, or use matchAll and String.match, which do not leave the cursor where you can trip over it. This page sidesteps the problem entirely by compiling a private working copy for every run, so the regex you are looking at is never the one being mutated.

What can PCRE do that JavaScript regex cannot?+

Several things, and knowing which ones saves a lot of confused debugging when you copy a pattern across. There are no atomic groups: PCRE's (?>…) commits to a match and refuses to backtrack into it, which is the standard hand-fix for catastrophic backtracking, and JavaScript's only substitute is a lookahead wrapping a capture. There are no possessive quantifiers (a*+) for the same reason. There is no recursion (?R) and there are no subroutine calls, so balanced brackets are out of reach. There are no inline modifiers: (?i) mid-pattern is a syntax error, because a JavaScript flag applies to the whole regex or not at all. There are no conditionals, no \A \z \Z anchors, no POSIX classes such as [[:alpha:]], and no comment mode. What JavaScript does have that many engines do not is variable-length lookbehind.

Does \d mean the same thing everywhere?+

No, and the difference is a real source of security bugs. In JavaScript \d is exactly [0-9], always, with or without the u flag. In .NET and Python 3, \d matches any character with the Unicode decimal-digit property by default, which includes Arabic-Indic digits ٠١٢, Devanagari digits, and fullwidth 012. A validator written in Python that accepts ٢٠٢٦ and a JavaScript one that rejects it will disagree about the same input, which is precisely the sort of mismatch that gets exploited. The same caution applies to \w, which is [A-Za-z0-9_] here and so does not consider “é” a word character, and to \b, which is defined in terms of \w and therefore lands in the middle of “naïve”. If you want Unicode semantics in JavaScript you have to ask for them explicitly: \p{Nd} with the u flag for digits, \p{L} for letters.

When should I stop using a regular expression?+

When the thing you are matching can nest. Regular expressions in the formal sense cannot count, and although backreferences push JavaScript's engine past a strictly regular language, it still cannot match balanced structure. That rules out HTML, XML, JSON, source code, and any bracket language, not because the pattern is hard to write but because no pattern exists. Use DOMParser, JSON.parse, or a real parser. Three softer signals point the same way: the pattern has grown past a line or two and nobody can read it, it needs a comment to explain each clause, or it keeps acquiring special cases for inputs it was supposed to already handle. A regular expression is at its best finding and validating flat, well-shaped fragments (a date, a hex colour, a log prefix), and at its worst pretending to be a grammar.

What do $1, $&, $` and $' mean in a replacement?+

They are the substitution patterns String.replace understands in the replacement string. $1 through $99 insert the text of that numbered capture group, and $<name> does the same for a named one. $& inserts the whole match. $` inserts everything before the match and $' inserts everything after it; both are surprisingly useful for wrapping and surprisingly easy to trigger by accident, since a backslash does not escape them. $$ inserts a single literal dollar sign. Two rules catch people out: a reference to a group that does not exist is left in the output as plain text rather than throwing, and $12 means group 12 if the pattern has twelve groups and group 1 followed by the character “2” if it does not. This page expands all of them in the preview and flags any reference it could not resolve.

More focused tools, ready when you are.

Explore the growing collection for calculations, documents, writing, and everyday work.

Browse all tools