Regex patterns/Text & extraction

Repeated word

A backreference catches the the kind of typo a spellchecker walks straight past.

/\b(\w+)\s+\1\b/gi
Open it in Rex

The problem

Find places where the same word appears twice in a row — the commonest proofreading miss.

How it reads

word edge1\wrepeat\srepeatsame as 1word edge

Follow the line from left to right — every path you can trace is a string this pattern matches.

  1. \b(\w+)\s+\1\bIn order:
  2. \bA word boundary — the edge between a word character and anything else
  3. (\w+)Capture group 1:
  4. \w+A word character (letter, digit or underscore), one or more times, as many as possible
  5. \s+Whitespace, one or more times, as many as possible
  6. \1The same text that group 1 captured earlier
  7. \bA word boundary — the edge between a word character and anything else

Matches

  • This is is a test
  • The the quick brown fox
  • that that

Does not match

  • A fine sentence
  • a banana
  • is island

Where it bites

  • A backreference matches the *text* the group captured, not the pattern — which is why \1 here means "that same word again", not "another word".
  • The trailing \b is essential: without it, "is island" matches, because the second word only has to start the same way.
  • The i flag makes "The the" a hit. Drop it if capitalisation should count as a difference.

More text & extraction patterns