Regular Expression Cheat Sheet

A regular expression is a tiny program for matching text. This reference covers the syntax shared by JavaScript, Python, Java, .NET, Go and PCRE, points out where they differ, and ends with patterns you can copy.

Characters and classes

.        any character except a line break (with the s flag: any character)
\d \D    digit / not a digit
\w \W    word character [A-Za-z0-9_] / not one
\s \S    whitespace / not whitespace
\b \B    word boundary / not a word boundary
\t \n \r tab, line feed, carriage return
\.       a literal dot (escape any of . * + ? ^ $ { } ( ) [ ] | \ /)
[abc]    one of a, b or c
[^abc]   any character except a, b or c
[a-z0-9] ranges inside a class
\p{L}    any Unicode letter (JavaScript needs the u or v flag)

Inside square brackets most metacharacters lose their meaning: [.+*] matches a literal dot, plus or star. Only ], \, ^ (first position) and - (between characters) are special, and a - placed first or last is literal.

In JavaScript, \d and \w are ASCII-only even with the u flag, whereas Python 3 and .NET match Unicode digits and letters by default. Use [0-9] when you mean ASCII digits in a portable pattern.

Anchors and quantifiers

^  $        start / end of input (of each line with the m flag)
a*          zero or more        a*?   lazy: as few as possible
a+          one or more         a+?   lazy
a?          zero or one         a??   lazy
a{3}        exactly three
a{2,}       two or more
a{2,5}      two to five         a{2,5}? lazy
a++  a*+    possessive (PCRE, Java, Python 3.11+; not JavaScript)

Quantifiers are greedy by default: they match as much as possible, then give characters back if the rest of the pattern fails. <.+> on <b>bold</b> matches the whole string; the lazy <.+?> matches <b> only. A negated class such as <[^>]+> is usually better than either, because it cannot run past the closing bracket.

Python, Ruby and PCRE also offer \A and \z (\Z in Python) for the absolute start and end of the input, unaffected by multiline mode. JavaScript has no equivalent; use ^ and $ without the m flag.

Groups, alternation and backreferences

(abc)          capturing group, numbered from 1 by its opening parenthesis
(?:abc)        non-capturing group
(?<year>\d{4}) named group (JavaScript, .NET, Java, PCRE)
(?P<year>...)  named group in Python, referenced later as (?P=year)
\1  \k<year>   backreference to group 1 / to a named group
a|b            alternation: a or b
(?>...)        atomic group (PCRE, Java, .NET, Python 3.11+; not JavaScript)

Alternation has the lowest precedence of all operators. ^cat|dog$ means “cat at the start, or dog at the end”; write ^(?:cat|dog)$ to anchor both. Inside a replacement string, JavaScript refers to groups as $1 and $<year>, and to the whole match as $&; Python uses \1 and \g<year>.

Lookarounds

Lookarounds test what comes before or after the current position without including it in the match:

\d+(?=px)     digits followed by "px"         (positive lookahead)
\d+(?!px)     digits not followed by "px"     (negative lookahead)
(?<=\$)\d+    digits preceded by "$"          (positive lookbehind)
(?<!-)\d+     digits not preceded by "-"      (negative lookbehind)

Lookaheads are supported everywhere except Go’s RE2-based regexp package, which deliberately omits lookarounds and backreferences to guarantee linear-time matching. Lookbehind is supported in modern JavaScript engines, .NET and PCRE; Python and Java require the lookbehind to have a bounded or fixed length, while JavaScript and .NET allow any length.

Stacked lookaheads at the start of a pattern express “must contain” rules: ^(?=.*[a-z])(?=.*[A-Z])(?=.*\d).{12,}$ requires a lower-case letter, an upper-case letter and a digit, in any order, with at least 12 characters.

Flags

JavaScript flags, written after the closing slash or passed as the second argument to new RegExp:

  • g global — find all matches (required by matchAll and replaceAll with a regex)
  • i ignore case
  • m multiline — ^ and $ match at line breaks
  • s dotAll — . also matches line breaks
  • u Unicode — code-point matching, \p{...} and \u{...} escapes, stricter syntax
  • v Unicode sets — everything u does plus set operations in classes such as [\p{L}--[a-z]]
  • y sticky — match only at lastIndex
  • d indices — report start and end positions of each group

Other flavours have equivalents with different spellings — Python’s re.IGNORECASE, re.MULTILINE, re.DOTALL and re.VERBOSE (free-spacing mode, x in PCRE) — and many accept inline flags such as (?i) at the start of the pattern.

One JavaScript trap: a regex with the g or y flag remembers lastIndex between calls to test() or exec(), so calling test() twice on the same string can return true and then false.

Common patterns

These are pragmatic patterns for validation and extraction, not complete parsers. Each is anchored; remove ^ and $ to search inside longer text.

ISO date        ^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$
24-hour time    ^([01]\d|2[0-3]):[0-5]\d$
UUID            ^[0-9a-f]{8}-[0-9a-f]{4}-[1-8][0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$   (i flag)
IPv4 address    ^((25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)\.){3}(25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)$
Hex colour      ^#([0-9a-f]{3}|[0-9a-f]{6})$   (i flag)
Semantic ver.   ^\d+\.\d+\.\d+(-[0-9A-Za-z.-]+)?(\+[0-9A-Za-z.-]+)?$
Slug            ^[a-z0-9]+(-[a-z0-9]+)*$
Decimal number  ^-?\d+(\.\d+)?$
Email (rough)   ^[^\s@]+@[^\s@]+\.[^\s@]+$
Trailing spaces [ \t]+$       (m flag)

The date pattern accepts 2026-02-31; check calendar validity in code. For email, a loose pattern plus a confirmation message beats any attempt at the full RFC 5322 grammar.

Flavour differences in one place

  • JavaScript: no atomic groups or possessive quantifiers, no \A/\z; named groups as (?<n>...); lookbehind of any length; Unicode classes only with u or v.
  • Python (re): named groups as (?P<n>...); fixed-width lookbehind only; atomic groups and possessive quantifiers since 3.11; re.fullmatch checks the whole string without anchors. The third-party regex module adds variable-length lookbehind and more Unicode features.
  • Java: close to PCRE; backslashes must be doubled inside string literals ("\\d+"), and String.matches implicitly anchors the whole input.
  • .NET: the richest flavour, with variable-length lookbehind, balancing groups and a RegexOptions.NonBacktracking mode for untrusted input.
  • Go (regexp): RE2 syntax, linear-time guarantee, no lookarounds or backreferences.
  • PCRE (PHP, grep -P, nginx): possessive quantifiers, atomic groups, recursion and (*VERB) controls.

A pattern that uses only classes, anchors, quantifiers, groups and alternation will behave the same in all of them. Each feature beyond that is worth checking against the engine that will actually run it.

Reading an unfamiliar pattern

Faced with something like ^(?:[a-z0-9-]+\.)+[a-z]{2,}$, read it from the outside in. The anchors say the pattern must cover the entire input. Inside, a non-capturing group repeated one or more times matches a label of letters, digits and hyphens followed by a dot. The pattern ends with two or more letters. Put together: a host name such as api.example.com. Naming each piece like this, in order, is faster than trying to simulate the engine in your head, and it is exactly what an explainer tool automates.

Keeping regexes fast and maintainable

  • Anchor patterns whenever you can; an anchored pattern fails fast on non-matching input.
  • Prefer specific classes ([^,]*) to .*, which forces the engine to backtrack.
  • Avoid nested quantifiers over overlapping characters, such as (\w+\s?)+$. On a long input that almost matches, backtracking engines can take exponential time — the cause of several real outages.
  • Build long patterns from named pieces in code, or use verbose mode with comments, so the next reader does not have to decode a 200-character line.

To see what an unfamiliar pattern does, paste it into the regex explainer, which describes each part in English and draws a railroad diagram.

Frequently asked questions

What is the difference between greedy and lazy quantifiers?

Greedy quantifiers (, +, {n,m}) match as much as possible and backtrack when needed. Lazy versions (?, +?, {n,m}?) match as little as possible and extend only when the rest of the pattern requires it.

How do I write a named group in JavaScript and Python?

JavaScript uses (?<name>…) and refers back with \k<name>. Python uses (?P<name>…) and (?P=name).

Does JavaScript support lookbehind?

Yes, in all current major browsers and Node.js, and unlike Python it allows lookbehinds of variable length.

Why does my regex with the g flag alternate between true and false?

With g or y, RegExp.prototype.test updates lastIndex, so the next call starts searching where the previous match ended. Create a new regex or reset lastIndex to 0.

Can a regex validate an email address?

Only roughly. The full address grammar is too complex for a practical regex; use a simple pattern and confirm the address by sending a message.

Related