What is a regular expression, and when should you not use one?
A regular expression (regex) is a pattern describing a set of strings, used for searching, matching, extracting and replacing text. It is one of the highest-leverage tools available and one of the most frequently misapplied.
The basic elements: literal characters match themselves; character classes ([a-z], \d, \w) match sets; quantifiers (*, +, ?, {2,5}) specify repetition; anchors (^, $, \b) match positions; groups ( ) capture portions for extraction; and alternation | matches one of several options.
Where regex is the right tool: validating simple formats, extracting fields from consistently structured text, search and replace across a codebase, log parsing, and tokenising.
Where it is the wrong tool:
Parsing HTML or XML. This is the canonical example. HTML is not a regular language — it permits arbitrary nesting, which regular expressions fundamentally cannot express. Any regex handling nested tags will fail on valid input. Use a parser.
Parsing JSON, or any recursive structure, for the same reason.
Email validation. The fully correct RFC-compliant pattern is notoriously enormous and still does not tell you whether the address exists. Check for an @ with something either side, then send a confirmation email, which is the only real validation.
Anything with a proper parser available — dates, URLs, CSV. CSV in particular defeats naive regex through quoted fields containing commas and newlines.
Two practical dangers:
Catastrophic backtracking. Nested quantifiers such as (a+)+b can take exponential time on certain inputs. This is a genuine denial-of-service vector — ReDoS — and has taken down production services. Avoid nested quantifiers on user input, and prefer engines with linear-time guarantees.
Unreadability. A complex regex is write-only. Use verbose mode with comments, named capture groups, and break it into pieces.
Test against real data, including the cases you did not anticipate.