Regex Cheat Sheet
Every regular expression token, with a worked example for each — and notes on where JavaScript, Python, Java, Go, PCRE and .NET disagree.
Frequently asked questions
What is the difference between greedy and lazy?
A greedy quantifier takes as much as it can and then gives characters back until the rest of the pattern fits; a lazy one takes as little as possible and adds more only when forced. So against <a><b>, the pattern <.*> matches the whole string while <.*?> matches just <a>. Being specific — <[^>]*> — is usually better than either, because it cannot backtrack at all.
When should I use a non-capturing group?
Whenever you need grouping but not the captured text — which is most of the time. (?:…) keeps your group numbers meaningful, avoids storing text you will not read, and makes the pattern marginally faster. Reserve numbered groups for things you actually extract.
Why does \d match more than 0-9 sometimes?
In Python, \d matches any Unicode decimal digit by default, including Arabic-Indic and Devanagari numerals — pass re.ASCII to restrict it. JavaScript, Java and Go all keep \d as ASCII 0-9 unless you ask otherwise. If a pattern validates input, this difference matters.
What does the "u" flag actually change?
It makes JavaScript treat the pattern as a sequence of Unicode code points rather than UTF-16 code units. That fixes matching against characters outside the Basic Multilingual Plane — emoji and many CJK characters — and it is required for \p{…} property escapes. It also makes several previously-tolerated escapes into syntax errors, which is a feature.
Found a mistake on this page? Tell me — a page that is confidently wrong is worse than no page.