All posts

Why Email Validation Regex Fails, and What to Use Instead

··9 min read

Email validation regex fails in two directions at once: popular patterns reject real addresses (plus tags, apostrophes, long TLDs, internationalised domains) and accept broken ones (double dots, 65-character local parts). Some also backtrack catastrophically. Use a regex only for rough shape, then check length limits, normalise the domain, and confirm MX records.

We did not want to argue this from first principles, so we measured it.

The test: five regexes, 23 addresses

None of the five patterns we tested got more than 15 of 23 cases right. The best performers were the WHATWG pattern browsers use and a simple "no spaces, one @, one dot" pattern, and they failed on different things.

The five patterns are the ones you find over and over in tutorials, Stack Overflow answers and form libraries:

Label Pattern
A, unanchored /\S+@\S+\.\S+/
B, anchored non-space /^[^\s@]+@[^\s@]+\.[^\s@]+$/
C, "classic" /^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$/
D, old tutorial /^\w+([\.-]?\w+)*@\w+([\.-]?\w+)*(\.\w{2,3})+$/
E, WHATWG HTML the pattern from the HTML standard's email input definition

We ran each against 23 addresses in Node.js 22 on 1 October 2026. "Expected" is what a sensible signup form should do: accept anything the standards allow and real mail systems deliver, reject anything malformed. The raw output is in our research notes.

Address Expected A B C D E
jane+newsletter@example.com accept yes yes yes no yes
o'brien@example.ie accept yes yes no no yes
jane@example.photography accept yes yes yes no yes
user@xn--mnchen-3ya.de accept yes yes yes no yes
user@münchen.de accept yes yes no no no
"john doe"@example.com accept (RFC-valid) yes no no no no
john..doe@example.com reject yes yes yes no yes
.john@example.com reject yes yes yes no yes
john@-example.com reject yes yes yes no no
john doe@example.com reject yes no no no no
john@@example.com reject yes no no no no
65-character local part reject yes yes yes yes yes
john@example reject no no no no yes
Correct out of 23 13 15 13 14 15

The table shows the interesting rows; the full 23 also included ordinary addresses that every pattern handled.

Three things stand out.

Pattern A accepts john doe@example.com because it is not anchored. Without ^ and $, .test() succeeds if any substring matches, and doe@example.com does. This is the single most common bug in copied email regexes.

Pattern D rejects plus tags, apostrophes and every TLD longer than three letters. .photography, .museum and punycode domains all fail. It was written when TLDs were mostly two or three characters, and that assumption has been wrong for over a decade.

Every pattern accepted a 65-character local part. None of them check length, and length is where the standards are actually strict.

What the standards actually allow

The local part (before the @) is far more permissive than most regexes assume, and the length limits are stricter than most regexes check.

RFC 5322 section 3.2.3 defines atext, the characters allowed in an unquoted local part: letters, digits and ! # $ % & ' * + - / = ? ^ _ ` { | } ~. Dots are allowed between runs of those characters but not at the start, the end, or twice in a row. A quoted local part such as "john doe" can contain almost anything.

RFC 5321 section 4.5.3.1 sets the size limits:

Part Limit Source text
Local part 64 octets "The maximum total length of a user name or other local-part is 64 octets."
Domain 255 octets "The maximum total length of a domain name or number is 255 octets."
Path 256 octets Includes the < and > around the address, so the address itself tops out at 254

Each DNS label in the domain (the parts between dots) is limited to 63 characters and cannot start or end with a hyphen, which is why the WHATWG pattern has that {0,61} in it.

Then there is internationalised email. RFC 6531 extends SMTP so that both the local part and domain can contain UTF-8, which means 用户@例子.广告 is a legitimate address on a server that supports it. Only patterns A and B accepted it, and only because they accept nearly anything.

Why the WHATWG pattern is the sensible default

It is the pattern every browser already applies to <input type="email">, it is linear-time, and its authors are explicit about which parts of RFC 5322 they chose to ignore.

The HTML standard calls its definition "a willful violation of RFC 5322", because RFC 5322 is "simultaneously too strict (before the '@' character), too vague (after the '@' character), and too lax (allowing comments, whitespace characters, and quoted strings in manners unfamiliar to most users) to be of practical use here."

That is the right trade-off for a signup form. Quoted local parts and comments are legal but vanishingly rare, and an address like "john doe"@example.com will break plenty of downstream systems even if your form accepts it.

Its weaknesses, from our table: it allows leading, trailing and doubled dots, it accepts john@example (a valid hostname, but not something deliverable on the public internet), it rejects Unicode domains unless you convert them to punycode first, and it does no length checking.

The regex that takes five seconds to say no

Pattern D, the old tutorial regex, backtracks exponentially: rejecting a 33-character string took 4.8 seconds on an Apple M2.

The problem is \w+([\.-]?\w+)*. Because the separator is optional, a run of letters can be split between \w+ and the repeated group in an exponential number of ways. When the string fails to match at the end, the engine tries every one of them.

We fed it "a".repeat(n) + "!" and timed .test():

const tutorial = /^\w+([\.-]?\w+)*@\w+([\.-]?\w+)*(\.\w{2,3})+$/;
for (const n of [24, 26, 28, 30, 32]) {
  const input = "a".repeat(n) + "!";
  const t0 = process.hrtime.bigint();
  tutorial.test(input);
  const ms = Number(process.hrtime.bigint() - t0) / 1e6;
  console.log(`${n + 1} chars: ${ms.toFixed(0)} ms`);
}
Input length Time to reject (Node 22, M2)
25 characters 91 ms
29 characters 297 ms
31 characters 1,169 ms
33 characters 4,772 ms

Two runs gave the same numbers to within a few milliseconds. Every two extra characters roughly quadruples the time. JavaScript regexes run on the main thread, so in a Node server one crafted signup request stalls every other request for the duration. This is the class of bug OWASP documents as ReDoS, and one of the "evil regex" examples OWASP lists is itself an email-validation pattern.

The WHATWG pattern, by contrast, rejected a 10,000-character string in well under a millisecond in the same test.

What to use instead

Split the address on the last @, check lengths, convert the domain to ASCII, check the local part and each domain label separately, and then let DNS answer the question a regex cannot.

This function runs in Node 18 or later with no dependencies. We tested it against the same cases as above.

import { domainToASCII } from "node:url";

const ATEXT = "[A-Za-z0-9!#$%&'*+/=?^_`{|}~-]";
const LOCAL_PART = new RegExp(`^${ATEXT}+(\\.${ATEXT}+)*$`);
const DNS_LABEL = /^(?!-)[A-Za-z0-9-]{1,63}(?<!-)$/;

export function checkEmailSyntax(input) {
  const email = String(input).trim();
  const at = email.lastIndexOf("@");
  if (at < 1 || at === email.length - 1) {
    return { ok: false, reason: "needs text either side of @" };
  }

  const local = email.slice(0, at);
  const domain = domainToASCII(email.slice(at + 1).toLowerCase());

  if (!domain) return { ok: false, reason: "domain is not a valid hostname" };
  if (local.length > 64) return { ok: false, reason: "local part over 64 octets" };
  if (local.length + 1 + domain.length > 254) return { ok: false, reason: "address over 254 octets" };
  if (!LOCAL_PART.test(local)) return { ok: false, reason: "local part has characters we do not accept" };

  const labels = domain.split(".");
  if (labels.length < 2 || !labels.every((label) => DNS_LABEL.test(label))) {
    return { ok: false, reason: "domain labels are malformed" };
  }

  return { ok: true, email: `${local}@${domain}` };
}

What it does with our test cases:

Input Result
Jane.Doe@Example.COM ok, Jane.Doe@example.com (trimmed, domain lower-cased)
user@münchen.de ok, user@xn--mnchen-3ya.de
o'brien@example.ie, jane+news@example.com ok
john..doe@, .john@, john@-example.com rejected
65-character local part rejected, "local part over 64 octets"
用户@例子.广告 rejected

A few deliberate choices are worth calling out.

It lower-cases the domain but not the local part. RFC 5321 section 2.4 says the local part "MUST BE treated as case sensitive", even though almost no real mailbox provider treats it that way. Store it as typed; compare case-insensitively if you need to deduplicate.

It rejects quoted local parts and non-ASCII local parts. Both are legal. Both are rare, and both break enough CRMs, ESPs and CSV exports that most products are better off refusing them with a clear message. If your audience includes users with internationalised addresses, relax LOCAL_PART and make sure your sending provider supports SMTPUTF8 first.

There is no nested quantifier that can backtrack. The dot is mandatory between runs of ATEXT, so there is only one way to match any given string. It rejected a 100,000-character input in 0.05 ms.

It returns a reason. "Please enter a valid email" is useless to someone whose real address was rejected. Telling them what failed lets them correct a typo or contact you.

A regex is the first check, not the answer

A syntax check catches typing mistakes. It cannot tell you that the domain receives mail or that the mailbox exists.

asdkjhaslkdjh@gmail.com passes every pattern above. So does jane@gmial.com, which is a typo of a real domain and the subject of our post on catching typos at signup. The layers that actually answer "will this deliver" are:

  1. Syntax, as above, which proves the string is shaped like an address.
  2. DNS, which proves the domain can receive mail. That means MX records, the implicit-MX fallback in RFC 5321 section 5.1, and the null MX record from RFC 7505 that a domain publishes to say it never accepts mail. We cover the code for that in validating email in JavaScript and Node.js and validating email in Python.
  3. SMTP, which asks the recipient's server whether the mailbox exists. This needs outbound port 25, which most cloud hosts block.

The comparison of what each layer catches is in syntax check vs MX check vs SMTP verification.

The practical takeaway

  • If you only add one thing, anchor your regex. An unanchored pattern accepts strings with spaces and double @ signs.
  • Prefer the WHATWG pattern or a split-and-check function like the one above over anything with \w+([.-]?\w+)* in it. Test any pattern you inherit with a long string of letters followed by !.
  • Check lengths explicitly: 64 for the local part, 254 for the whole address.
  • Convert Unicode domains to punycode before matching, or you will reject real users at münchen.de.
  • Do not stop at syntax. A DNS lookup is cheap and catches dead domains, and you can try a single address in our free email verifier to see what the deeper checks return.

Common questions

What is the best regex for email validation?

The most defensible choice is the pattern published in the WHATWG HTML standard, because it is what every browser uses for input type=email and it is safe from catastrophic backtracking. It still accepts some undeliverable addresses and rejects internationalised ones, so treat it as a shape check, not proof the address works.

Can a regex tell me if an email address exists?

No. A regex only looks at the characters. Whether the domain receives mail needs a DNS lookup for MX records, and whether the mailbox exists needs an SMTP conversation with the recipient's server. A perfectly formed address can still bounce.

Why does my email regex reject addresses with a plus sign?

Because many tutorial patterns only allow letters, digits, dots, underscores and hyphens before the @. RFC 5322 also allows characters such as + ' ! # $ % & * / = ? ^ ` { | } ~ in the local part, and the plus sign in particular is used by Gmail and Microsoft 365 (Exchange Online), among others, for sub-addressing.

Can an email regex crash my server?

Yes, if it contains nested quantifiers that backtrack. In our test a widely copied pattern took 4.8 seconds to reject a 33-character string in Node.js, and every two extra characters roughly quadrupled the time. Node runs that on the main thread, so one request can stall every other request.

Verify unlimited addresses for $29.99/month

Real SMTP mailbox checks. No credits, no per-email fees.

Get Started