Regex patterns/Text & extraction

HTML tag

Finds tags in a blob of markup — useful for stripping, dangerous for parsing.

/</?([a-z][\w-]*)[^>]*>/gi
Open it in Rex

The problem

Find the tags in a chunk of markup so you can strip them or count them.

How it reads

</skip1[a-z][\w-]repeatskip[^>]repeatskip>

Follow the line from left to right — every path you can trace is a string this pattern matches.

  1. </?([a-z][\w-]*)[^>]*>In order:
  2. <The character "<"
  3. /?The character "/", optionally (zero or one time)
  4. ([a-z][\w-]*)Capture group 1:
  5. [a-z][\w-]*In order:
  6. [a-z]Any one of: "a" to "z"
  7. [\w-]*Any one of: a word character (letter, digit or underscore) or "-", any number of times, including none, as many as possible
  8. [^>]*Any character except ">", any number of times, including none, as many as possible
  9. >The character ">"

Matches

  • <p class="lead">
  • </strong>
  • <br/>

Does not match

  • plain text
  • a < b and c > d

Where it bites

  • HTML is not a regular language, so no regex can parse it correctly. Attributes containing > break this one, and comments and CDATA break every version of it.
  • For anything beyond a quick strip or a count, use DOMParser in the browser or a real parser on the server.
  • [^>]* is what keeps the pattern linear; a .*? here would be slower and no more correct.

More text & extraction patterns