Regex Cheat Sheet

Every regular expression token, with a worked example for each — and notes on where JavaScript, Python, Java, Go, PCRE and .NET disagree.

Written to be read in one pass and then skimmed forever afterwards. Every row has an example, because a token you cannot picture is a token you will misuse.

Character classes

TokenMatchesExample
.Any character except a line break (any character at all with the s flag)a.c → abc, a-c
[abc]Any one of the listed charactersgr[ae]y → gray, grey
[^abc]Any character not listed[^aeiou] → a consonant, a digit, a space
[a-z]A range[a-fA-F0-9] → a hex digit
\d \DA digit / not a digit\d{4} → 2026
\w \WLetter, digit or underscore / the opposite\w+ → user_42
\s \SWhitespace / not whitespace\s{2,} → two or more spaces
\p{L} \P{L}A character with / without a Unicode property. Needs the u flag\p{L}+ → letters in any script

Quantifiers

TokenRepeatsExample
*Zero or moreab*c → ac, abc, abbbc
+One or moreab+c → abc but not ac
?Zero or onecolou?r → color, colour
{3}Exactly three\d{3} → 123
{2,}Two or more\d{2,} → 12, 12345
{2,5}Between two and five\d{2,5} → 12 to 12345
*? +? {2,5}?The same, but lazy — as few as possible<.*?> → <a> rather than the whole line
*+ ++Possessive — never gives characters back. Not in JavaScript, Python or Go\d++ → Java, PCRE and .NET only

Anchors and boundaries

TokenMatches
^Start of the string — or of each line with the m flag
$End of the string — or of each line with the m flag
\bA word boundary: the edge between \w and \W. Matches a position, consuming nothing
\BAny position that is not a word boundary
\A \z \ZAbsolute start / absolute end / end before a trailing newline. Not in JavaScript

\b is worth understanding properly, because it is the fix for the most common regex bug there is: cat matches inside concatenate, while \bcat\b does not.

Groups and references

TokenPurposeExample
(…)Capture, numbered from 1 by opening bracket(\d{4})-(\d{2}) → year, month
(?:…)Group without capturing(?:ab)+ → repeats ab
(?<name>…)Named capture(?<year>\d{4}) → match.groups.year
\1 \2Backreference — the text a group already matched\b(\w+)\s+\1\b → a doubled word
\k<name>Backreference to a named group(?<q>["']).*?\k<q>
a|bAlternation. First match wins, so order matterscat|category never matches category fully

Lookaround

Lookaround checks what is next to the current position without consuming it, so the match itself excludes what was tested.

TokenMeansExample
(?=…)Followed by\d+(?= USD) → the number in 50 USD
(?!…)Not followed by\d+(?! USD) → numbers not in dollars
(?<=…)Preceded by(?<=\$)\d+ → the number in $50
(?<!…)Not preceded by(?<!\$)\d+ → numbers not after a dollar sign

Lookarounds combine well for validation: ^(?=.*[a-z])(?=.*[A-Z])(?=.*\d).{8,}$ means “at least eight characters, containing at least one lowercase letter, one uppercase letter and one digit” — each lookahead tests independently from the start.

Flags

FlagEffectCaution
gFind every match, not just the firstA global regex keeps lastIndex between calls — reusing one across test() calls gives alternating results
iIgnore caseCase folding for non-ASCII varies by engine
m^ and $ match at line boundariesDoes not make . match newlines — that is s
s. matches line breaks tooCalled DOTALL or Singleline elsewhere
u / vUnicode mode; required for \p{…}Turns some previously-tolerated escapes into errors
ySticky — match only at lastIndexUseful for tokenisers, surprising elsewhere

Escaping

These characters are special and need a backslash to match literally:

. ^ $ * + ? ( ) [ ] { } | \ /

In JavaScript you can escape a runtime string safely with str.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"), or RegExp.escape() where it is available. Python has re.escape(), Java has Pattern.quote(), and Go has regexp.QuoteMeta(). Build a pattern from user input any other way and you have an injection bug.

Performance

Regex performance is almost entirely about backtracking. Four rules cover most of it:

  1. Never nest quantifiers. (a+)+ gives the engine exponentially many ways to fail. This is the classic denial-of-service pattern, and it takes about twenty characters of input to hang a thread.
  2. Prefer a negated class to a lazy dot. [^"]* cannot backtrack; .*? can.
  3. Anchor where you can. ^ stops the engine retrying at every position.
  4. Order alternatives cheapest-first, and make sure the first one that can match is the one you want.

The regex tester abandons matching after about 800 milliseconds and tells you — which is a fast way to find out that a pattern backtracks badly before it reaches production.

Where the engines differ

FeatureJSPythonJavaGoPCRE.NET
LookaheadYesYesYesNoYesYes
LookbehindYes, variableFixed widthBoundedNoFixed widthYes, variable
BackreferencesYesYesYesNoYesYes
Named groups(?<n>)(?P<n>)(?<n>)(?P<n>)BothBoth
Possessive / atomicNoNoYesNoYesYes
Inline flags (?i:…)NoYesYesYesYesYes
Linear-time guaranteeNoNoNoYesNoOptional

Go’s omissions are deliberate. RE2 drops the features that require backtracking and gets a linear-time guarantee in exchange, which is why it is the right choice when the pattern or the input comes from a user.

Frequently asked questions

What is the difference between greedy and lazy?

A greedy quantifier takes as much as it can and then gives characters back until the rest of the pattern fits; a lazy one takes as little as possible and adds more only when forced. So against <a><b>, the pattern <.*> matches the whole string while <.*?> matches just <a>. Being specific — <[^>]*> — is usually better than either, because it cannot backtrack at all.

When should I use a non-capturing group?

Whenever you need grouping but not the captured text — which is most of the time. (?:…) keeps your group numbers meaningful, avoids storing text you will not read, and makes the pattern marginally faster. Reserve numbered groups for things you actually extract.

Why does \d match more than 0-9 sometimes?

In Python, \d matches any Unicode decimal digit by default, including Arabic-Indic and Devanagari numerals — pass re.ASCII to restrict it. JavaScript, Java and Go all keep \d as ASCII 0-9 unless you ask otherwise. If a pattern validates input, this difference matters.

What does the "u" flag actually change?

It makes JavaScript treat the pattern as a sequence of Unicode code points rather than UTF-16 code units. That fixes matching against characters outside the Basic Multilingual Plane — emoji and many CJK characters — and it is required for \p{…} property escapes. It also makes several previously-tolerated escapes into syntax errors, which is a feature.

Found a mistake on this page? Tell me — a page that is confidently wrong is worse than no page.