Regex patterns/Text & extraction
Words in a text
Counting what a human calls a word, apostrophes included.
/\b[\w']+\b/gThe problem
Count or list the words in a block of text, treating don’t as one word rather than two.
How it reads
Follow the line from left to right — every path you can trace is a string this pattern matches.
\b[\w']+\bIn order:\bA word boundary — the edge between a word character and anything else[\w']+Any one of: a word character (letter, digit or underscore) or "'", one or more times, as many as possible\bA word boundary — the edge between a word character and anything else
Matches
- doesn't
- hello world
- one
Does not match
- ...
- !!!
Where it bites
- \w is [A-Za-z0-9_], so digits and underscores count as words and accented letters do not. For anything multilingual, Intl.Segmenter is the right tool.
- The apostrophe has to be added explicitly; \b[\w]+\b splits don't into two.
- Hyphenated words still split in two — add - to the class if they should not.