
How to Choose Fonts for Multilingual Websites?
- Sajjad
- Typography
- 01 Oct, 2026
The English site looks great in a display serif the brand team picked. Then the Vietnamese and Polish translations go live, and the headlines turn into a patchwork: most letters render in the brand font, but every ơ, ư, ą, and ł drops to a system fallback with a different weight and height. On the Japanese pages, the brand font does not cover kanji at all, so the whole page renders in whatever the visitor's operating system chooses, and it looks different on Windows, macOS, and Android. None of this showed up in design reviews, because every mockup was in English.
This article explains how to evaluate a font's language coverage, how to choose families that work across scripts, how to build per-script font stacks with unicode-range and :lang(), how to harmonize mixed-script text, and how to keep multilingual font loading fast.
What "Multilingual Support" Actually Means for a Font
A font supports a language when it contains every character that language needs and renders them correctly. That involves three layers:
- Script coverage. The writing system: Latin, Cyrillic, Greek, Arabic, Devanagari, Han, and so on. A Latin-only font cannot render Russian, no matter what.
- Character set within a script. "Latin" ranges from basic ASCII to hundreds of accented letters. English needs about 100 characters. Vietnamese needs dozens of extra letters with stacked diacritics, like ặ and ở. Polish needs ą, ę, ł, ń, ś, ź, ż. A font labeled "Latin" may only cover Western European languages.
- Language-specific behavior. Some languages share characters but expect different shapes. Romanian uses ș and ț with a comma below, not a cedilla. Bulgarian Cyrillic has distinct letterforms from Russian Cyrillic. Serbian italics differ from Russian italics. Good fonts handle this through the OpenType
locl(localized forms) feature, which the browser applies based on thelangattribute.
When a font lacks a character, the browser does not show a blank. It falls back to the next font in your stack, or a system font, for that one character. That is why missing coverage looks like a mismatched letter in the middle of a word rather than an obvious error.
Step 1: List Your Languages and Their Scripts
Before looking at any fonts, write down every language the site supports now and is likely to support within a couple of years. Group them by script and note anything unusual:
| Language | Script | Notes |
|---|---|---|
| English | Latin | Basic Latin |
| German | Latin | ä, ö, ü, ß, and capital ẞ; long compound words |
| Polish | Latin | Latin Extended-A characters |
| Vietnamese | Latin | Stacked diacritics; needs extra line height |
| Turkish | Latin | Dotted and dotless i; case conversion is language-specific |
| Russian | Cyrillic | Full Cyrillic block |
| Arabic | Arabic | Right-to-left, joined letterforms |
| Hindi | Devanagari | Complex shaping, conjuncts, tall marks |
| Japanese | Han + Kana | Thousands of glyphs; large files |
This table drives every decision that follows. A Latin-plus-Cyrillic site can usually use one family. A site that spans Latin, Arabic, Devanagari, and Japanese needs a strategy for combining families.
Step 2: Verify Coverage Before You Commit
Do not trust marketing pages that say a font "supports 100+ languages." Check the actual file.
On Google Fonts, filter by language on the browse page, and check the family's glyph view for the specific characters you need. The language filter is a good first pass.
For any font file, inspect the character map with fontTools:
pip install fonttools brotli
# check_coverage.py
from fontTools.ttLib import TTFont
font = TTFont("BrandSerif-Regular.woff2")
cmap = font.getBestCmap()
samples = {
"vi": "ăâđêôơưạảấầẩẫậắằẳẵặẹẻẽếềểễệỉịọỏốồổỗộớờởỡợụủứừửữựỳỵỷỹ",
"pl": "ąćęłńóśźżĄĆĘŁŃÓŚŹŻ",
"ro": "ăâîșțĂÂÎȘȚ",
"tr": "çğıİöşüÇĞÖŞÜ",
"de": "äöüßẞÄÖÜ",
}
for lang, chars in samples.items():
missing = [c for c in chars if ord(c) not in cmap]
print(f"{lang}: {'OK' if not missing else 'missing ' + ''.join(missing)}")
Run it against every weight and style you plan to load. Italic and bold files sometimes have smaller character sets than the regular weight, especially in older or cheaper families.
Then test rendering, not just presence. Build a test page with real sentences in each language, set lang correctly on each block, and look at it in Chrome, Safari, and Firefox. Check that diacritics do not collide with the line above, that Romanian shows comma-below letters, and that bold and italic do not fall back.
Step 3: Choose a Family Strategy
There are three realistic approaches.
One Family with Broad Coverage
Some families cover many scripts in a consistent design. This is the simplest option when your languages fit within its coverage:
- Noto (Google): the broadest coverage of any family, with separate Noto Sans and Noto Serif families for nearly every script, all designed to harmonize.
- IBM Plex: Latin, Cyrillic, Greek, plus dedicated Plex Sans Arabic, Hebrew, Devanagari, Thai, Japanese, and Korean families.
- Inter: Latin (including Vietnamese), Cyrillic, and Greek, with excellent screen rendering.
- Source Sans 3 and Source Han Sans: Adobe's Latin family and its matching CJK superfamily.
A Brand Font Plus Matched Script Partners
If the brand font is Latin-only, pair it with families designed for other scripts that have similar structure: comparable stroke contrast, x-height relative to cap height, and weight range. A geometric Latin sans pairs better with a geometric Arabic like a Kufi style than with a calligraphic Naskh. The general principles of font pairing still apply, just across scripts instead of within one.
System Fonts for Some Scripts
For scripts with huge character sets, especially Chinese, Japanese, and Korean, using the operating system's fonts is often the pragmatic choice. Every modern OS ships high-quality CJK fonts, and you avoid multi-megabyte downloads. The trade-off is less control over appearance. A system font stack per script keeps this predictable.
Step 4: Build Per-Script Stacks with unicode-range
The unicode-range descriptor tells the browser which characters a font face covers. The browser only downloads a face if the page contains characters in its range. That lets you register several fonts under one family name, each covering a different script:
/* Latin brand font */
@font-face {
font-family: "Site Sans";
src: url("/fonts/brand-sans-latin.woff2") format("woff2");
font-weight: 100 900;
font-display: swap;
unicode-range: U+0000-00FF, U+0131, U+0152-0153, U+02BB-02BC, U+02C6, U+02DA,
U+02DC, U+2000-206F, U+20AC, U+2122, U+2212, U+FEFF, U+FFFD;
}
/* Latin Extended and Vietnamese from the same brand font */
@font-face {
font-family: "Site Sans";
src: url("/fonts/brand-sans-latin-ext.woff2") format("woff2");
font-weight: 100 900;
font-display: swap;
unicode-range: U+0100-024F, U+1E00-1EFF, U+20AB;
}
/* Cyrillic from a matched partner font */
@font-face {
font-family: "Site Sans";
src: url("/fonts/partner-sans-cyrillic.woff2") format("woff2");
font-weight: 100 900;
font-display: swap;
unicode-range: U+0400-04FF, U+0500-052F, U+2116;
}
/* Arabic from a matched partner font */
@font-face {
font-family: "Site Sans";
src: url("/fonts/partner-sans-arabic.woff2") format("woff2");
font-weight: 100 900;
font-display: swap;
unicode-range: U+0600-06FF, U+0750-077F, U+08A0-08FF, U+FB50-FDFF, U+FE70-FEFF;
}
body {
font-family: "Site Sans", system-ui, sans-serif;
}
Now a Russian page downloads only the Latin and Cyrillic files, an English page downloads only the Latin file, and the CSS stays simple. Splitting a font into these files is done with pyftsubset; the process is in what font subsetting is.
Step 5: Use lang and :lang() for Language-Specific Rules
unicode-range handles scripts. It cannot handle languages that share a script but need different treatment, and it cannot choose between Chinese, Japanese, and Korean, which share thousands of Han code points with different preferred glyph shapes. For that, set the lang attribute and use the :lang() selector:
<html lang="en">
...
<blockquote lang="ja">吾輩は猫である。名前はまだ無い。</blockquote>
</html>
:lang(ja) {
font-family: "Site Sans", "Hiragino Sans", "Noto Sans JP", "Yu Gothic", sans-serif;
line-height: 1.8;
}
:lang(zh-Hans) {
font-family: "Site Sans", "PingFang SC", "Noto Sans SC", "Microsoft YaHei", sans-serif;
line-height: 1.8;
}
:lang(ko) {
font-family: "Site Sans", "Apple SD Gothic Neo", "Noto Sans KR", "Malgun Gothic", sans-serif;
word-break: keep-all;
}
:lang(vi) {
line-height: 1.7;
}
Putting "Site Sans" first means Latin characters inside Japanese text, such as product names, still render in the brand font, while kanji and kana fall through to the CJK fonts.
The lang attribute does more than select fonts. It enables the right locl glyph variants, correct hyphenation with hyphens: auto, correct quotation marks with the q element, correct case conversion with text-transform (Turkish i and İ), and correct pronunciation by screen readers. Set it on html for every page and on any element containing a different language. In Next.js, that typically means setting it from the locale segment in the root layout, as shown in how to create a multi-language website with Next.js.
Step 6: Harmonize Mixed-Script Text
When two scripts appear side by side, such as English product names in Arabic text or Latin numerals in Hindi, differences in apparent size are obvious. Two tools help.
font-size-adjust scales fallback fonts so their x-heights match the first font's. It is supported in all major browsers:
body {
font-family: "Site Sans", system-ui, sans-serif;
font-size-adjust: ex-height 0.52;
}
Measure your primary font's x-height ratio (x-height divided by font size) and use that value. Fallback faces get scaled to match, which reduces both visual mismatch and layout shift.
Script-specific sizing and line height. Some scripts need more vertical room. Vietnamese stacks diacritics above capitals. Devanagari and Bengali have marks above the headline and below the baseline. Thai has stacked vowels and tone marks. Arabic generally reads best slightly larger than Latin at the same nominal size. Set these with :lang() rules rather than one global value:
| Script / language | Typical body line height | Notes |
|---|---|---|
| Latin (English) | 1.5–1.6 | WCAG 1.4.12 baseline |
| Vietnamese | 1.65–1.75 | Stacked diacritics |
| Arabic | 1.7–1.9 | Tall ascenders, deep descenders |
| Devanagari, Bengali | 1.7–1.9 | Marks above and below the headline |
| Thai | 1.7–1.9 | Stacked vowel and tone marks |
| Chinese, Japanese | 1.7–1.9 | Dense glyphs need more air |
Do not apply Latin habits to other scripts. Negative letter spacing breaks Arabic joining and can make Devanagari conjuncts collide. Synthetic italics on CJK just slant square glyphs. Detailed handling of those scripts is in how to display Bengali, Hindi, and other non-Latin scripts and right-to-left typography in CSS.
Step 7: Plan for Text Expansion
Translations change length. German and Finnish text is often 20–35% longer than English; Chinese and Japanese are often shorter but taller. Font choice affects how well a layout absorbs that:
- Avoid condensed display fonts for UI labels that must fit in fixed widths. They look fine in English and overflow in German.
- Never fix button or tab widths. Let them grow with content.
- Test headings with long compounds like "Datenschutzerklärung" and enable
hyphens: autowith a correctlangso the browser can break them. - Use
text-wrap: balanceon headings so expanded translations wrap evenly.
Step 8: Keep Loading Fast
Multilingual sites can easily load several megabytes of fonts if you are not careful:
- Self-host and subset by script, so each page downloads only what it uses.
- Preload only the primary Latin file (or the primary script of each locale), never every subset.
- Use variable fonts where possible to cover all weights in one file per script.
- Let system fonts handle CJK unless the brand requires otherwise; if you do serve CJK web fonts, use a service or build process that slices them into many small
unicode-rangechunks.
With next/font, request only the subsets each font needs:
import { Noto_Sans, Noto_Sans_Arabic } from "next/font/google";
export const notoSans = Noto_Sans({
subsets: ["latin", "latin-ext", "vietnamese", "cyrillic"],
variable: "--font-noto-sans",
display: "swap",
});
export const notoSansArabic = Noto_Sans_Arabic({
subsets: ["arabic"],
variable: "--font-noto-arabic",
display: "swap",
preload: false,
});
Setting preload: false on secondary scripts stops Next.js from preloading files that most pages will never use.
Multilingual Fonts FAQ
Check the font's character map for every character the language needs, using a tool like fontTools or the glyph view on Google Fonts, then test real sentences in the browser with the correct lang attribute. Check every weight and style, since some have smaller character sets.
Noto offers the broadest script coverage with a consistent design, and IBM Plex covers many major scripts with dedicated families. If your brand font is Latin-only, pair it with partner fonts designed for each additional script.
The font is missing those characters, so the browser renders them with a fallback font. Choose a font with full coverage for that language or add a matched partner font for the missing range.
Yes. The lang attribute enables language-specific glyph forms, hyphenation, correct case conversion, correct quotation marks, and accurate screen reader pronunciation. It is also how CSS can choose the right Chinese, Japanese, or Korean font for shared characters.
Often system fonts are the better choice, because every major operating system ships good CJK fonts and full web font files are several megabytes. If you need a specific CJK typeface, serve it in many small unicode-range slices so pages download only the characters they use.
The browser checks the characters on the page against each font face's unicode-range and downloads only the faces that are needed. An English page never downloads the Cyrillic or Arabic file registered under the same family name.
Conclusion
Choosing fonts for a multilingual site starts with an honest list of languages and scripts, not with the brand font. Verify coverage in the actual files for every weight, decide whether one broad family, a brand font with matched partners, or system fonts suit each script, and remember that languages within one script can still need different glyphs and spacing.
Then implement it cleanly. Register script-specific files under one family name with unicode-range, set lang everywhere and use :lang() for language-specific stacks and line heights, harmonize sizes with font-size-adjust, and design layouts that absorb text expansion. Done this way, each visitor downloads only the fonts their language needs, and every page looks deliberate in every language you publish.
Here are some useful references for going deeper on multilingual typography:
- MDN Web Docs: unicode-range — syntax and behavior of the descriptor that powers per-script font loading.
- W3C Internationalization: Declaring language in HTML — how and why to set the lang attribute.
- Google Fonts Knowledge: Choosing type — guidance on selecting fonts, including language support considerations.
- MDN Web Docs: font-size-adjust — matching x-heights across fonts and fallbacks.
- Google Fonts: Noto — the Noto project's families for nearly every script.


