What Is an IDN Homograph Attack?
An IDN homograph attack registers a lookalike domain using Unicode characters that resemble Latin letters. How it works, how browsers react, and defenses.
An IDN homograph attack is a kind of domain spoofing in which an attacker registers a domain name containing Unicode characters that look the same as, or nearly the same as, ordinary Latin letters. For example, the Cyrillic “а” looks just like the Latin “a” but is a different character. Swap one for the other in a familiar brand name and you get a domain that reads correctly to a human but resolves to a server the attacker controls. That domain can then be used for phishing pages, malware downloads, or credential theft.
Internationalized domain names and Punycode
The DNS was designed around a small ASCII character set: letters, digits, and hyphens. Internationalized domain names (IDNs) let people register domains in their own scripts, including Cyrillic, Greek, Arabic, Chinese, and many others. They’re important for making the web usable in non-English-speaking regions.
Because the underlying DNS still works in ASCII, IDNs are converted into an ASCII-compatible form called Punycode before lookup. A Punycode label starts with the prefix xn--, followed by an encoding of the Unicode characters. Browsers display the friendly Unicode version in the address bar, while resolvers and servers see the xn-- version.
You can see this conversion in JavaScript. The WHATWG URL API normalizes hostnames to their ASCII form:
// The first letter below is Cyrillic "а" (U+0430), not Latin "a"
const url = new URL("https://аpple.example/");
console.log(url.hostname); // "xn--pple-43d.example"
The visible text and the actual destination have separated, and that gap is what homograph attacks exploit.
How the attack works
Unicode contains many confusables, which are characters from different scripts that look alike in common fonts. Some well-known pairs:
| Looks like | Actual character | Script |
|---|---|---|
| a | а (U+0430) | Cyrillic |
| e | е (U+0435) | Cyrillic |
| o | о (U+043E) | Cyrillic |
| p | р (U+0440) | Cyrillic |
| o | ο (U+03BF) | Greek |
| l | ӏ (U+04CF) | Cyrillic |
An attack generally follows these steps:
- Pick a target, such as a bank, an email provider, or a popular SaaS login page.
- Build a lookalike by replacing one or more Latin letters with confusables. Some brand names can be spelled entirely with characters from another script, which produces a “whole-script” homograph with no mixing at all.
- Register the domain and get a valid TLS certificate for it. A padlock in the address bar only proves the connection is encrypted to that domain, as explained in how HTTPS works. It says nothing about whether it’s the domain you meant to visit.
- Distribute the link by email, text message, social media, or ads, and host a cloned login page.
It’s a close relative of typosquatting. Typosquatting relies on users mistyping, while homograph attacks rely on users misreading, and a careful reader may not be able to tell the difference at all.
How browsers defend against it
Browsers know about this problem and use display rules to decide whether to show a hostname in Unicode or fall back to raw Punycode. The details vary by browser, but the common checks include:
- Mixed-script detection. A label that combines scripts that don’t normally appear together, such as Latin and Cyrillic, is shown as Punycode.
- Whole-script confusable checks. A label written entirely in one non-Latin script but made only of letters that look Latin can be flagged, especially when it resembles a well-known domain.
- Lookalike comparison. Some browsers compare the hostname’s “skeleton” (its visual shape after mapping confusables) against a list of popular domains and fall back to Punycode when there’s a match.
- Unicode security guidance. Many of these rules come from the Unicode Consortium’s security mechanisms specification, which defines confusable mappings and restriction levels for identifiers.
Domain registries also play a part. Many restrict which scripts can be used under a given top-level domain and reject labels that mix scripts.
None of these defenses is complete. Display rules use heuristics, fonts differ, and other places a link might appear, such as email clients, chat apps, and PDF viewers, may not apply the same checks.
Defenses that actually hold up
The strongest defenses don’t depend on a person spotting the difference at all.
- Phishing-resistant authentication. Passkeys and other WebAuthn credentials are bound to the exact origin they were created on. A passkey for the real domain won’t be offered on a lookalike, however convincing the page looks.
- Password managers. Autofill matches the exact domain stored with the credential. If your manager doesn’t offer to fill a login form, treat that as a warning sign.
- Certificate monitoring. Every publicly trusted certificate is logged, so brands can watch Certificate Transparency logs for newly issued certificates on lookalike domains and act before a campaign starts.
- Defensive registration. Organizations sometimes register the most obvious confusable versions of their primary domain.
- Email authentication. SPF, DKIM, and DMARC stop attackers from spoofing your exact domain, which pushes them toward lookalikes. Those lookalikes can then be caught by mail filters that check sender domains for confusables.
Homographs beyond domain names
The same trick shows up wherever people judge text by how it looks. Lookalike characters have been used in package names on public registries, in usernames on social platforms, and in source code, where a variable named with a Cyrillic letter can look identical to a legitimate one during review. Tooling that flags non-ASCII identifiers or mixed scripts, such as linters, registry policies, and code-review warnings, applies the browser’s approach to these other areas.
The takeaway
An IDN homograph attack takes advantage of the fact that many Unicode characters from different scripts look identical, so a domain can read as a trusted brand while pointing somewhere else entirely. Browsers reduce the risk by showing suspicious hostnames as xn-- Punycode, but those rules are heuristics and don’t cover every place a link can appear. Reliable protection comes from defenses that check the exact domain rather than how it looks: passkeys, password-manager autofill, and certificate monitoring for the domains that impersonate yours.
Tagged
Keep reading
Chisato · · 4 min read DNS over HTTPS vs DNS over TLS
DoH tunnels DNS queries inside HTTPS on port 443; DoT wraps them in TLS on a dedicated port 853. Both encrypt lookups — here's how they differ.
Chisato · · 4 min read What Is DNS Tunneling? Hiding Data in DNS Queries
DNS tunneling encodes data inside DNS queries and responses to smuggle traffic past firewalls, since DNS is almost always allowed through unfiltered.
Chisato · · 4 min read IPv4 vs IPv6: What's Actually Different
IPv4's 32-bit address space is exhausted; IPv6 fixes that with 128-bit addresses plus routing and header changes. Here's what differs in practice.