Regex patterns/Text & extraction

Words in a text

Counting what a human calls a word, apostrophes included.

/\b[\w']+\b/g
Open it in Rex

The problem

Count or list the words in a block of text, treating don’t as one word rather than two.

How it reads

word edge[\w']repeatword edge

Follow the line from left to right — every path you can trace is a string this pattern matches.

  1. \b[\w']+\bIn order:
  2. \bA word boundary — the edge between a word character and anything else
  3. [\w']+Any one of: a word character (letter, digit or underscore) or "'", one or more times, as many as possible
  4. \bA word boundary — the edge between a word character and anything else

Matches

  • doesn't
  • hello world
  • one

Does not match

  • ...
  • !!!

Where it bites

  • \w is [A-Za-z0-9_], so digits and underscores count as words and accented letters do not. For anything multilingual, Intl.Segmenter is the right tool.
  • The apostrophe has to be added explicitly; \b[\w]+\b splits don't into two.
  • Hyphenated words still split in two — add - to the class if they should not.

More text & extraction patterns