Writing Secure Regex Rules

Several Traffic Server configuration files and plugins accept regular expressions for matching incoming requests. When the regex result drives a security or routing decision (rule selection, ACL allow/deny, signature exclusion, parent selection, SNI routing, and so on), the way the operator writes the regex matters: an unanchored or partially-anchored pattern can match more inputs than the operator intended, including inputs crafted by clients to fire a rule that was meant to apply to something else.

This page documents:

  • the regex matching contract used at security-sensitive call sites,

  • the input subject each call site matches against,

  • common pitfalls that produce over-matching patterns, and

  • recommended pattern shapes for each site.

The matching contract

Traffic Server uses PCRE2 for regex matching. Patterns are compiled once at config load and reused per request. By default, a successful match means the pattern matched some substring of the input subject — possibly the whole subject, possibly a prefix, possibly a fragment in the middle. The matcher does not, by default, require the pattern to consume the entire subject.

This is the standard PCRE behavior. It has security implications when the subject is influenced by a client (a request URL, a host header, an SNI value, a Referer header, etc.). An operator who writes the pattern cdn\.example\.com as a host regex thinking “match this exact host” will find the rule firing on client-supplied hosts like cdn.example.com.other.org, because the substring cdn.example.com appears at position 0 of the longer subject.

The fix at the operator level is to anchor patterns explicitly so the pattern says exactly what it should match:

  • ^pattern$ — match exactly the full input

  • ^pattern — match anything starting with pattern

  • pattern$ — match anything ending with pattern

  • ^.*pattern.*$ — match anything containing pattern (when substring matching is the actual intent)

Anchored patterns produce predictable behavior regardless of the matcher’s defaults at any given site. Unanchored or single-anchored patterns may behave one way today and another way after future changes; the only durable approach is to write patterns whose intent is unambiguous from the pattern itself.

Subjects by call site

The “subject” — the string the regex is matched against — varies by site. Knowing the subject is essential for writing correct patterns, because the same regex against different subjects gives different results.

remap.config — regex_map

Subject: request host (no scheme, port, or path).

For URL http://cdn.example.com:8080/path, the regex sees the subject cdn.example.com.

remap.config — map_with_referer

Subject: the value of the Referer HTTP header (header value only, not including the Referer: field name or trailing CRLF).

For header Referer: https://www.partner.com/page, the regex sees the subject https://www.partner.com/page.

parent.config — url_regex

Subject: the full request URL including scheme, host, and path.

For URL http://example.com/news/politics/today, the regex sees the subject http://example.com/news/politics/today.

cache.config — host_regex

Subject: request host (no scheme, port, or path) — same as regex_map.

cache.config — url_regex

Subject: the full request URL — same as parent.config’s url_regex.

splitdns.config — url_regex

Subject: the full request URL — same as parent.config’s url_regex.

url_sig — excl_regex

Subject: the full request URL, sliced before the first ? or #.

For URL http://host/path?query, the regex sees the subject http://host/path.

maxmind_acl — country regex

Subject: host + "/" + path (no scheme, no query string).

For URL http://example.com/file.txt, the regex sees the subject example.com/file.txt.

geoip_acl — country regex

Subject: the request URL path returned by TSUrlPathGet, which does not include the leading /.

For URL http://example.com/song.mp3, the regex sees the subject song.mp3. For URL http://example.com/foo/song.mp3, the regex sees the subject foo/song.mp3.

tls_bridge — SNI routing

Subject: the SNI value from the inbound TLS ClientHello.

Patterns at this site are start-anchored at compile time by the plugin, so prefix injection is blocked. Operators should still add the end-anchor $ to block trailing-content matches.

uri_signing — cdniuc regex match

Note

The “matching contract” section at the top of this document describes Traffic Server’s general regex behavior, which is implemented with PCRE2. This call site is the exception: uri_signing uses the GNU regex.h library’s re_match, which is start-anchored by definition (the pattern always matches starting at offset 0 of the subject). A leading ^ in an issuer pattern is therefore implicit, and only the trailing anchor $ controls whether trailing content is allowed — the opposite of PCRE2’s default substring match.

Subject: the normalized request URI (scheme, authority, path, and query — produced by the plugin’s URI-normalization step).

Unlike the other sites in this document, the regex pattern at this site is not operator-controlled: it is the value of an inbound JWT’s cdniuc claim, in the form regex:<pattern>, and is chosen by the token issuer. The ATS operator is the verifier, not the author of the pattern. An unanchored issuer pattern that names a directory prefix (for example https://media\.example\.com/preview/) will accept any request URI that starts with that prefix — including URIs the issuer did not intend the token to authorize. Operators who deploy uri_signing should require their token issuers to fully anchor every regex: container with ^...$ and confirm the issuer rejects unanchored patterns at token-mint time.

The CDNI URI Signing draft itself does not explicitly require any anchoring — it states only that the URI must match the regex, leaving “match” open to interpretation. Traffic Server performs start-anchoring (the GNU re_match API used internally is start-anchored by definition) but does not enforce end-anchoring; other CDNI relying parties may interpret the contract differently. Fully anchored issuer patterns are the one shape every relying party agrees on, so they remove cross-implementation ambiguity in addition to bounding the issuer’s intended scope.

Common pitfalls

Bare-token patterns without anchors

A pattern like \.pdf without any anchor matches anywhere in the input. Against the subject http://host/protected.pdf.alternate, the substring .pdf appears at the expected position, so the rule fires — even though the URL does not actually end in .pdf.

Recommended: .*\.pdf$ (suffix anchor) for “URLs ending in .pdf”, or ^http://[^?#]*\.pdf$ (full anchor) for full-URL matching against a subject that includes the scheme.

DNS-label boundaries

A pattern like cdn\.example\.com matches anywhere in the host subject. Against cdn.example.com.other.org, the substring matches at position 0 and the rule fires — but the operator probably meant “the exact host cdn.example.com.”

Recommended: ^cdn\.example\.com$ for an exact-host match, or ^.*\.example\.com$ for “any subdomain of example.com.” Note that simply adding the start-anchor ^ is not enough: ^cdn\.example\.com (start-anchored only) still matches cdn.example.com.other.org because the start matches and the end is unanchored.

Forgetting that the subject excludes the scheme

For regex_map, cache.config host_regex, maxmind_acl, and geoip_acl, the subject does not include the URL scheme (http://). Patterns that try to match a leading http:// will never fire at these sites; check the Subjects by call site section above before writing the pattern.

Forgetting that geoip_acl strips the leading slash

For geoip_acl, the subject is the URL path returned by TSUrlPathGet, which strips the leading /. A pattern like /songs/.*\.mp3 will never fire against this subject; use ^songs/.*\.mp3$ instead.

Suffix-only anchoring may not express the full operator intent

A pattern like \.pdf$ does match subjects that end in .pdf — PCRE scans forward through the subject and the suffix anchor is satisfied at end-of-input. The pattern works.

What the pattern does not do is constrain the rest of the subject: it matches http://example.com/file.pdf and cdn.example.com/file.pdf and a.pdf, regardless of the rest of the input. At sites whose subject is path-only or host-only (for example, geoip_acl or cache.config host_regex), this can match more contexts than the operator had in mind.

For clarity, prefer patterns that document the full operator intent. ^https?://.*\.pdf$ is more verbose than \.pdf$ but makes the expected subject shape (an HTTP/HTTPS URL ending in .pdf) self-documenting and resistant to surprise if the same pattern is copied to a different site whose subject shape differs.

Why this matters

Operator-written regex rules can become security boundaries when they gate behavior such as access control, signature verification, upstream selection, or DNS routing. A regex that is permissive in ways the operator did not intend can let traffic through that the operator meant to block, route to an unintended upstream, or skip a verification step that should have applied.

Auditing existing regex rules to confirm they match exactly what the operator intends — and no more — is a one-time cost that pays off in predictable rule behavior across upgrades, configuration changes, and unexpected client inputs.