Regex patterns/Text & extraction
HTML tag
Finds tags in a blob of markup — useful for stripping, dangerous for parsing.
/</?([a-z][\w-]*)[^>]*>/giThe problem
Find the tags in a chunk of markup so you can strip them or count them.
How it reads
Follow the line from left to right — every path you can trace is a string this pattern matches.
</?([a-z][\w-]*)[^>]*>In order:<The character "<"/?The character "/", optionally (zero or one time)([a-z][\w-]*)Capture group 1:[a-z][\w-]*In order:[a-z]Any one of: "a" to "z"[\w-]*Any one of: a word character (letter, digit or underscore) or "-", any number of times, including none, as many as possible[^>]*Any character except ">", any number of times, including none, as many as possible>The character ">"
Matches
- <p class="lead">
- </strong>
- <br/>
Does not match
- plain text
- a < b and c > d
Where it bites
- HTML is not a regular language, so no regex can parse it correctly. Attributes containing > break this one, and comments and CDATA break every version of it.
- For anything beyond a quick strip or a count, use DOMParser in the browser or a real parser on the server.
- [^>]* is what keeps the pattern linear; a .*? here would be slower and no more correct.