Internationalized Domain Names and Homograph Risks
How Unicode support in domain names enables visually deceptive registrations that evade casual detection.
Core Concept
Internationalized Domain Names (IDNs) allow non-ASCII characters in domain names, encoded internally using Punycode, supporting genuine multilingual internet use.
The same capability that enables legitimate non-Latin domains also allows homograph attacks, where visually near-identical characters from different scripts spoof a trusted domain.
How Homograph Attacks Work
Certain characters across different scripts render nearly or entirely identically to the human eye.
- Cyrillic 'а' visually identical to Latin 'a' in many fonts
- Greek characters resembling Latin letters in some typefaces
- Mixed-script domains exploiting inconsistent browser rendering rules
Punycode Encoding
An IDN domain is internally represented in Punycode, prefixed with xn--, revealing the true encoded form even when the display looks identical to a trusted brand.
Browser Defenses Against Homographs
Modern browsers implement mitigations to reduce homograph attack effectiveness.
- Displaying Punycode instead of Unicode for suspicious mixed-script domains
- Restricting which script combinations are permitted to render as Unicode
- Warning indicators for domains matching known homograph patterns
Detecting Homograph Registrations
Brand protection monitoring for homograph domains requires generating and checking the space of visually similar variants.
This involves mapping known homoglyph character substitutions across relevant scripts, then checking domain registries for any matching registrations.
Homoglyph Character Maps
Comprehensive homoglyph mapping across Cyrillic, Greek, and other scripts is the foundation of any serious IDN-based brand monitoring program.
Response Strategies
Once a homograph registration is identified, response options mirror those for standard typosquatting.
- Dispute resolution processes for confirmed bad-faith registrations
- Direct takedown requests to the registrar or host
- Proactive defensive registration of the most likely variants
Real-World Implementation
Homograph monitoring is a specialized but increasingly standard part of brand protection programs.
- Financial institutions actively monitoring for homograph phishing domains
- Browser vendors continuously refining mixed-script display policies
- Security researchers publishing updated homoglyph character databases
As Unicode domain support continues to expand internet accessibility globally, homograph-aware monitoring remains an essential companion practice for any recognizable brand.
Common Mistakes to Avoid
A few common mistakes weaken defenses against homograph domain attacks.
- Monitoring only Latin-script typosquatting variants while ignoring homoglyph substitutions.
- Assuming browser mitigations catch every possible homograph combination.
- Failing to defensively register the most visually deceptive variants ahead of time.
- Overlooking single-script IDN domains that can still deceive despite browser protections.
- Not maintaining an updated homoglyph character map across relevant scripts.
- Overlooking newer scripts and character sets as browser support and abuse patterns evolve.
- Assuming homograph monitoring is only relevant for large, globally recognized brands.
- Failing to test how a domain actually renders across different browsers and operating systems.
- Overlooking that some fonts render certain homoglyphs more convincingly than others.
- Assuming homograph risk is limited to well-known Cyrillic and Greek substitutions.
- Failing to test suspected homograph domains on actual mobile device browsers.
- Overlooking that some registrars have their own restrictions on mixed-script registrations.
Best Practices Checklist
These practices strengthen defenses against homograph-based domain abuse.
- Maintain a comprehensive, updated homoglyph character map across relevant scripts.
- Defensively register the most visually deceptive variants for high-value brand domains.
- Don't rely solely on browser-level mitigations to catch every possible variant.
- Monitor domain registries continuously for newly registered homograph matches.
- Combine homograph monitoring with standard typosquatting detection for full coverage.
- Update homoglyph coverage as new scripts and character combinations become exploitable.
- Extend homograph monitoring to any brand with meaningful international recognition.
- Test rendering across multiple browsers and operating systems for suspected homograph domains.
- Test suspected homograph domains across different fonts, since rendering convincingness varies.
- Expand homoglyph awareness beyond just the most commonly cited Cyrillic and Greek examples.
- Include mobile device browsers specifically when testing homograph domain rendering.
- Check registrar-specific mixed-script registration policies, which can offer an additional layer of protection.
Frequently Asked Questions
Frequently asked questions about internationalized domain names and homograph risks.
What is a homograph attack?
It's the use of visually near-identical characters from different Unicode scripts, such as Cyrillic or Greek, to spoof a trusted domain's appearance.
How is an IDN domain actually represented internally?
It's encoded in Punycode, prefixed with xn--, which reveals the true encoded form even when the display looks identical to a trusted brand.
Do browsers fully protect against homograph attacks?
They implement mitigations like displaying Punycode for suspicious mixed-script domains, but coverage varies and isn't a complete guarantee.
Is IDN support itself a bad idea given the risk?
No — it's a genuine accessibility feature enabling non-Latin domain names; homograph abuse is a side effect, not the purpose of the capability.
What's the most reliable defense against homograph domains?
Defensively registering the most visually deceptive variants ahead of time remains one of the most reliable preventive measures available.
Is homograph monitoring only necessary for large global brands?
No — any brand with meaningful recognition, even regionally, can be a target, so monitoring scope shouldn't be limited to only the largest companies.
Does a suspicious IDN domain render the same across all browsers?
Not necessarily — mixed-script display policies vary between browsers, so testing across multiple platforms gives a more complete picture.
Do new character sets create new homograph risks over time?
Yes — as Unicode support expands and new scripts become commonly used, homoglyph mapping needs ongoing updates to stay comprehensive.
Does font choice affect how convincing a homograph attack looks?
Yes — some fonts render certain character substitutions more convincingly than others, affecting how easily a user might be deceived.
Is homograph risk limited to Cyrillic and Greek scripts?
No — other scripts and character sets can also produce visually confusable characters worth including in a comprehensive monitoring program.
Do registrars restrict mixed-script domain registrations?
Some do implement their own restrictions, adding a layer of protection beyond what browsers alone provide.
Check for Homograph Variants
Run a domain intelligence lookup to check for suspicious IDN variants of a domain.
Launch Tool →