Type something to search...
What Are Internationalized Domain Names (IDN) and Punycode?

What Are Internationalized Domain Names (IDN) and Punycode?

A client in Germany registers münchen-bäckerei.de, a shop in Russia runs on a .рф domain, and a Japanese brand wants its name in kanji. They type their domains into a DNS dashboard and get back strings like xn--mnchen-bckerei-dib09a.de. Is something broken? No — that's Punycode, the encoding that lets DNS, which was built for a small set of ASCII characters, carry domain names written in any script.

This article explains what internationalized domain names (IDNs) are, how Punycode turns Unicode labels into ASCII xn-- labels, the difference between the IDNA2003 and IDNA2008 standards and why it still causes bugs, how to convert and resolve IDNs in code and on the command line, and how homograph attacks abuse look-alike characters. It assumes you're comfortable with how a fully qualified domain name is built out of labels.

What Is an Internationalized Domain Name?

An internationalized domain name (IDN) is a domain name that contains characters outside the traditional ASCII letters, digits, and hyphen. That includes accented Latin letters (café.example), entirely different scripts (пример.рф, 例え.jp), and right-to-left scripts such as Arabic and Hebrew.

IDNs can appear at any level:

  • Second-level and below: münchen.de, bücher.example, shop.café.example.
  • Top-level: internationalized country-code TLDs like .рф (Russia) and many generic TLDs in non-Latin scripts.

The DNS protocol itself was never changed to support Unicode. Instead, IDNs are handled entirely at the application layer: software converts the human-readable name into an ASCII-compatible form before sending it to DNS, and converts it back for display.

Why DNS Needs an Encoding

The hostname rules that every resolver, mail server, and registry relies on only allow letters a to z, digits, and hyphens, compared case-insensitively. Changing that would have required upgrading every DNS implementation on the internet at once — and DNS case-insensitivity only works for ASCII, so Ü and ü couldn't have been matched reliably anyway.

The solution, standardized as IDNA (Internationalizing Domain Names in Applications), was to define a reversible mapping from Unicode labels to ASCII labels. The ASCII form is called an A-label; the Unicode form is a U-label.

U-label (what people see)A-label (what DNS stores)
münchenxn--mnchen-3ya
bücherxn--bcher-kva
примерxn--e1afmkfd
рфxn--p1ai

Every A-label starts with the ACE prefix xn-- (ACE stands for ASCII Compatible Encoding). That prefix tells software, "this label is an encoded IDN — decode it for display." Registries forbid ordinary labels from having hyphens in the third and fourth positions precisely so that xn-- can never appear by accident.

How Punycode Works

Punycode, defined in RFC 3492, is the algorithm that produces the part after xn--. It's an instance of a general encoding called Bootstring, designed to be compact and unambiguous.

Encoding münchen works like this:

  1. Copy the basic characters. All ASCII characters in the label are copied in order: mnchen.
  2. Add a delimiter. If any basic characters were copied, a hyphen is appended: mnchen-.
  3. Encode the non-ASCII characters. Each non-ASCII code point is encoded as a sequence of ASCII letters and digits that describes which character it is and where to insert it, using variable-length integers and an adaptive bias. For ü (U+00FC) inserted at position 1, that produces 3ya.
  4. Add the prefix. The full A-label is xn--mnchen-3ya.

You can see the raw Punycode step in Python, which ships a punycode codec:

print("münchen".encode("punycode"))        # b'mnchen-3ya'
print(b"mnchen-3ya".decode("punycode"))    # münchen

The codec handles only the Punycode step, not the xn-- prefix or the normalization rules that IDNA adds on top, so don't use it directly for domain names. Use an IDNA-aware library instead, as shown below.

Punycode works on each label independently. In shop.münchen.de, only the middle label is encoded: shop.xn--mnchen-3ya.de.

IDNA2003 vs IDNA2008 vs UTS #46

Before encoding, a Unicode label has to be normalized and checked: uppercase mapped to lowercase, equivalent characters unified, and disallowed characters rejected. The rules for that step have changed over time, and the differences still matter.

  • IDNA2003 (RFC 3490, with Nameprep in RFC 3491) mapped a broad set of characters to others before encoding. For example, the German ß was mapped to ss, so faß.de became fass.de. It was tied to Unicode 3.2.
  • IDNA2008 (RFCs 5890 to 5894) is stricter and version-independent. It defines which code points are valid based on their properties, does not map characters (input must already be lowercase and normalized), and treats ß, the Greek final sigma ς, and a few joiner characters as distinct valid characters. faß.de becomes xn--fa-hia.de, a different domain from fass.de.
  • UTS #46 is a Unicode Consortium compatibility layer that applies IDNA2003-style mapping (like lowercasing) and then IDNA2008 validity rules. Browsers and the WHATWG URL Standard use it, which is why typing Bücher.example into a browser works even though strict IDNA2008 rejects the uppercase B.

The result is that two tools can produce different A-labels from the same input. Here's that difference in Python, using the standard library codec (IDNA2003) and the third-party idna package (IDNA2008, with optional UTS #46):

import idna

names = ["münchen.de", "bücher.example", "faß.de", "пример.рф", "Bücher.example"]

for name in names:
    builtin = name.encode("idna").decode("ascii")          # IDNA2003 (stdlib)
    try:
        strict = idna.encode(name).decode("ascii")         # IDNA2008
    except idna.IDNAError as exc:
        strict = f"rejected ({exc.__class__.__name__})"
    uts46 = idna.encode(name, uts46=True).decode("ascii")  # what browsers do
    print(f"{name:16} 2003={builtin:24} 2008={strict:30} uts46={uts46}")

print(idna.decode("xn--mnchen-3ya.de"))
münchen.de       2003=xn--mnchen-3ya.de        2008=xn--mnchen-3ya.de              uts46=xn--mnchen-3ya.de
bücher.example   2003=xn--bcher-kva.example    2008=xn--bcher-kva.example          uts46=xn--bcher-kva.example
faß.de           2003=fass.de                  2008=xn--fa-hia.de                  uts46=xn--fa-hia.de
пример.рф        2003=xn--e1afmkfd.xn--p1ai    2008=xn--e1afmkfd.xn--p1ai          uts46=xn--e1afmkfd.xn--p1ai
Bücher.example   2003=xn--bcher-kva.example    2008=rejected (InvalidCodepoint)    uts46=xn--bcher-kva.example
münchen.de

Most names encode identically under all three. The faß.de row shows the real divergence: the standard library sends you to fass.de, while modern browsers send you to xn--fa-hia.de — potentially a different owner. The last row shows strict IDNA2008 refusing uppercase input that UTS #46 lowercases first. Install the package with pip install idna. For new code that handles user-entered domains, use idna.encode(name, uts46=True) so your results match what browsers do.

Converting IDNs in Other Environments

Node.js and browsers

Node's url module and the WHATWG URL class both apply UTS #46 processing:

import { domainToASCII, domainToUnicode } from "node:url";

console.log(domainToASCII("münchen.de")); // xn--mnchen-3ya.de
console.log(domainToASCII("faß.de")); // xn--fa-hia.de
console.log(domainToUnicode("xn--e1afmkfd.xn--p1ai")); // пример.рф
console.log(new URL("https://bücher.example/").hostname); // xn--bcher-kva.example

domainToASCII produces the A-label form you'd use in a DNS query, and domainToUnicode converts back for display. The URL constructor does the same conversion automatically, which is why hostname always comes back in ASCII. Save the snippet as an .mjs file to run it with node.

Command line

The idn2 tool from GNU libidn2 converts names using IDNA2008 with UTS #46 mapping:

idn2 münchen.de
# xn--mnchen-3ya.de

idn2 --decode xn--e1afmkfd.xn--p1ai
# пример.рф

You can then query the A-label directly with dig:

dig +short xn--mnchen-3ya.de A

Recent dig builds compiled with libidn2 accept Unicode names and convert them for you, and display answers in Unicode when the terminal supports it. Use +noidnin and +noidnout to turn that off if you want to see exactly what goes over the wire. See how to use the dig command for more on its options.

IDNs in Zone Files and DNS Dashboards

DNS servers only ever see A-labels. In a zone file, you write the encoded form:

$ORIGIN xn--mnchen-3ya.de.
$TTL 3600

@       IN  A      192.0.2.10
www     IN  CNAME  xn--mnchen-3ya.de.
shop    IN  A      192.0.2.20
; café.xn--mnchen-3ya.de written as an A-label:
xn--caf-dma  IN  A  192.0.2.30

Many DNS dashboards accept Unicode input and convert it automatically, but some store whatever you type, so confirm the saved record with a lookup. Certificate authorities also issue certificates for the A-label form — a certificate for münchen.de lists xn--mnchen-3ya.de in its Subject Alternative Name field. If you're setting up redirects between the Unicode and ASCII-transliterated versions of a brand, redirecting a domain using DNS records covers the options.

Homograph Attacks

Unicode contains many characters that look identical or nearly identical to Latin letters. The Cyrillic а (U+0430) is visually indistinguishable from the Latin a in most fonts. An attacker can register a domain built from look-alike characters and present it as a familiar brand. This is an IDN homograph attack.

# Latin "apple" vs. the same word spelled entirely with Cyrillic look-alikes
print("apple.com".encode("idna"))                       # b'apple.com'
print("аррӏе.com".encode("idna"))  # b'xn--80ak6aa92e.com'

The second name renders like apple.com but is a completely different domain. In 2017, a researcher registered exactly that domain to demonstrate the problem, and several browsers at the time displayed it as apple.com in the address bar.

Defenses now operate at several levels:

  1. Browser display policies. Chrome, Firefox, Safari, and Edge show the Unicode form only when a name passes script-mixing and confusability checks; otherwise they display the xn-- form so the deception is visible.
  2. Registry policies. Many registries restrict which scripts can be used under a TLD and block registrations that are confusable with existing names, or bundle variants together.
  3. Your own monitoring. Brand owners often register or monitor obvious look-alike IDNs of their main domain, alongside typo domains.

Homograph domains are a phishing tool, and they're often paired with other techniques covered in DNS hijacking. If your application displays user-supplied domains — in email clients, chat apps, or admin panels — apply the same principle: show the A-label when the Unicode form mixes scripts.

Email and IDNs

IDNs in the domain part of an email address, like info@münchen.de, work with standard mail systems as long as clients convert the domain to its A-label before the DNS MX lookup. Fully internationalized addresses, where the local part before the @ is also non-ASCII, require SMTPUTF8 support (RFC 6531) on every server in the path. Support is widespread among large providers but far from universal, so a non-ASCII local part is still risky as a primary business address.

Practical Checklist

  • Register both forms if the brand needs it. A Unicode IDN and an ASCII transliteration (muenchen.de) are separate domains with separate owners.
  • Use UTS #46 processing in application code so your conversions match browsers.
  • Store and compare A-labels in databases and allowlists; convert to Unicode only for display.
  • Watch the deviation characters ß, ς, and zero-width joiners, which encode differently under IDNA2003 and IDNA2008.
  • Check the certificate covers the A-label of every IDN you serve over HTTPS.

IDN and Punycode FAQ

An IDN is a domain name containing characters outside basic ASCII, such as accented letters or non-Latin scripts. It is converted to an ASCII form for use in DNS.

The xn-- prefix marks a label as an encoded internationalized domain name. The characters after it are the Punycode encoding of the original Unicode label.

Punycode, defined in RFC 3492, is an algorithm that represents a Unicode string using only ASCII letters, digits, and hyphens. IDNA uses it to turn U-labels into A-labels.

No. DNS servers only store and compare ASCII A-labels. Browsers, mail clients, and other applications convert names to and from Unicode on either side of the lookup.

IDNA2003 mapped many characters before encoding, such as ß to ss, and was tied to Unicode 3.2. IDNA2008 is stricter, version-independent, and treats characters like ß as distinct, so some names encode differently.

In Python, use idna.encode(name, uts46=True) from the idna package. In Node.js, use domainToASCII from node:url. On the command line, the idn2 tool performs the same conversion.

It's a phishing technique that uses look-alike characters from other scripts, such as Cyrillic letters resembling Latin ones, to register a domain that appears identical to a trusted one.

Yes. Certificate authorities issue certificates for the A-label form of the domain, such as xn--mnchen-3ya.de, and browsers match it against the encoded hostname.

Yes. Many country-code and generic TLDs use non-Latin scripts, such as .рф for Russia, which is stored in the root zone as xn--p1ai.

Conclusion

Internationalized domain names let people use their own languages and scripts in web addresses, and Punycode makes that possible without changing DNS at all. Applications convert each Unicode label into an ASCII A-label prefixed with xn--, DNS resolves that ASCII name exactly as it would any other, and the application converts it back for display. The complexity lives in the normalization rules — IDNA2003, IDNA2008, and the UTS #46 compatibility layer browsers use — and in the security problems that look-alike characters create.

If you work with IDNs, keep three habits: use a UTS #46-aware library for conversions, store and compare the xn-- form internally, and treat any domain that mixes scripts with suspicion. With those in place, an IDN is just another domain name.

Here are some useful references for going deeper on IDNs and Punycode:

  1. RFC 3492: Punycode — the Bootstring-based encoding used for IDN A-labels.
  2. RFC 5890: Internationalized Domain Names for Applications (IDNA): Definitions and Document Framework — the entry point to the IDNA2008 standards.
  3. Unicode Technical Standard #46: Unicode IDNA Compatibility Processing — the mapping layer used by browsers and URL parsers.
  4. IANA: Repository of IDN Practices — the label generation tables registries publish for the scripts they allow.
  5. MDN Web Docs: URL: hostname property — how the URL API exposes hostnames, including encoded IDNs.
Tags :
Share :

Related Posts

What Is the Difference Between Authoritative and Recursive DNS Servers?

What Is the Difference Between Authoritative and Recursive DNS Servers?

When someone says "the DNS server," they could mean two completely different machines doing two completely different jobs. One kind of server holds t

Continue Reading
Can DNS settings affect website speed?

Can DNS settings affect website speed?

Yes, DNS settings can significantly affect the speed at which a website loads for its users. DNS, or Domain Name System, is often likened to the inte

Continue Reading
Can You Use a CNAME Record on the Root Domain?

Can You Use a CNAME Record on the Root Domain?

It is one of the most common DNS questions there is. Your hosting platform says "add a CNAME pointing to myapp.example-cdn.net," it works perfectly

Continue Reading