
What Is Font Subsetting, and How Does It Reduce File Size?
- Sajjad
- Typography
- 01 Oct, 2026
You download a font from a foundry for a client's English-language marketing site and drop the WOFF2 into the project. It weighs 310KB. The site's entire JavaScript bundle is smaller than that. Open the font in a glyph viewer and you find out why: it contains Latin, Latin Extended, Cyrillic, Greek, Vietnamese, hundreds of symbols, fractions, alternate numerals, and stylistic sets the design never uses. The site needs maybe 200 of the font's 2,500 glyphs.
Font subsetting is how you ship only the glyphs you need. This article explains what subsetting removes, how much it saves, how to subset with pyftsubset and glyphhanger, how unicode-range lets you split a font into on-demand pieces, and the mistakes that leave readers looking at empty boxes.
What Is Font Subsetting?
Font subsetting is the process of creating a smaller version of a font file that contains only a chosen set of characters (and the glyphs and features needed to render them). Everything else is stripped out.
A font file contains much more than glyph outlines:
| Font component | What it holds | Effect of subsetting |
|---|---|---|
Glyph outlines (glyf or CFF) | The vector shapes of every letter | Removed for unused characters; the biggest saving |
Character map (cmap) | Mapping from Unicode code points to glyphs | Trimmed to the kept characters |
Layout features (GSUB, GPOS) | Kerning, ligatures, alternates, small caps, mark positioning | Can be kept, trimmed, or dropped |
Hinting (fpgm, prep, cvt) | Instructions for rendering at small sizes on low-res screens | Can be dropped to save space |
Variation data (gvar, fvar) | Interpolation data for variable font axes | Kept, but scaled down with the outlines |
Names and metadata (name) | Font names, copyright, license text | Usually kept; required by most licenses |
The glyph outlines account for most of the file size, so removing scripts you do not use delivers most of the savings.
How Much Does It Save?
Savings depend on how many scripts the original font covers. Typical results for a well-equipped sans-serif in WOFF2:
| Version | Approximate size |
|---|---|
| Full font, all scripts | 250–350KB |
| Latin plus Latin Extended | 80–120KB |
| Basic Latin subset (Google's "latin" range) | 30–60KB |
| Basic Latin, no hinting | 20–45KB |
| Uppercase A–Z only (for a logo or drop cap) | 5–10KB |
A variable font with several axes stays larger, but the proportional savings are similar. Cutting 200KB from a font that renders your above-the-fold text can noticeably improve Largest Contentful Paint on mobile, and smaller files are far more likely to arrive inside the short window used by font-display: optional.
Subsetting by Script vs. by Content
There are two broad approaches:
- Script-based subsetting keeps entire Unicode ranges: all of Basic Latin, all of Latin Extended-A, all of Cyrillic. It is safe for user-generated and CMS content because any text in that script renders. This is what Google Fonts does.
- Content-based subsetting keeps only the exact characters found on your pages. It produces the smallest files but breaks the moment someone types a character you did not include. It suits fixed text: a logo wordmark, a navigation set in a display face, a single drop cap.
For body text on any site with editable content, use script-based subsetting. For decorative display faces used in a handful of fixed places, content-based subsetting is fine.
Subsetting with pyftsubset
pyftsubset is part of fontTools, the Python library maintained by the font engineering community and used by Google Fonts. Install it with Brotli support so it can write WOFF2:
pip install fonttools brotli
A Latin Subset for Body Text
pyftsubset SourceSerif4-Regular.ttf \
--unicodes="U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD" \
--layout-features="kern,liga,calt,ccmp,locl,mark,mkmk,onum,lnum,tnum,pnum,frac,sups,subs" \
--flavor=woff2 \
--output-file=source-serif-4-latin.woff2
The key options:
--unicodeslists the code points to keep. The range above is the same "latin" subset Google Fonts serves, covering English and most Western European languages, plus common punctuation, the euro sign, and the trademark symbol.--layout-featureslists the OpenType features to keep. Use"*"to keep them all. If you drop features likekernorliga, you lose kerning and ligatures in the subset. Keep any features your CSS turns on with font-feature-settings.--flavor=woff2writes compressed WOFF2 output.--output-filesets the output name.
Other Useful Options
# Subset to the exact characters in a text file
pyftsubset Display.ttf --text-file=wordmark.txt --flavor=woff2 --output-file=display-wordmark.woff2
# Subset to specific characters inline
pyftsubset Display.ttf --text="ABCDEFGHIJKLMNOPQRSTUVWXYZ" --flavor=woff2 --output-file=display-caps.woff2
# Drop hinting for extra savings on modern high-DPI screens
pyftsubset Inter.ttf --unicodes-file=latin.txt --no-hinting --desubroutinize --flavor=woff2 --output-file=inter-latin.woff2
--no-hinting removes TrueType hinting instructions. Modern macOS, iOS, and Android ignore hinting, and Windows renders most fonts acceptably without it at typical body sizes on high-density displays. If many readers use Windows on low-resolution monitors, test before dropping hints.
--desubroutinize expands CFF subroutines in OpenType CFF fonts. It makes the raw font larger but often compresses better in WOFF2, so the final file is smaller.
Subsetting a Variable Font
pyftsubset handles variable fonts the same way; the variation tables are subset along with the outlines. If you only need part of an axis range, first restrict the axis with fontTools' instancer, then subset:
fonttools varLib.instancer "Inter[opsz,wght].ttf" wght=400:700 opsz=14:32 -o inter-400-700.ttf
pyftsubset inter-400-700.ttf --unicodes-file=latin.txt --layout-features="*" --flavor=woff2 --output-file=inter-latin-400-700.woff2
Limiting the weight axis from 100–900 to 400–700 can cut the variation data substantially. For more on the trade-offs, see our comparison of variable fonts vs. static fonts.
Finding the Characters You Use with glyphhanger
glyphhanger, a Node tool by Zach Leatherman, crawls pages and reports the Unicode code points they use. It can also run the subset for you, using fontTools under the hood.
npm install -g glyphhanger
# Report the characters used across a site
glyphhanger https://example.com --spider --spider-limit=50
# Subset a font to those characters and output WOFF2
glyphhanger https://example.com --spider --subset=Display.ttf --formats=woff2
Use it to audit what characters your content contains, then round the result up to full script ranges for body fonts. Its output is a great starting point for display fonts used only in fixed places.
Splitting Subsets with unicode-range
You do not have to choose one subset. The unicode-range descriptor lets you declare several files for the same family, and the browser downloads each file only if the page contains characters in its range.
/* Latin: downloaded on almost every page */
@font-face {
font-family: "Source Serif 4";
src: url("/fonts/source-serif-4-latin.woff2") format("woff2");
font-weight: 400;
font-display: swap;
unicode-range: U+0000-00FF, U+0131, U+0152-0153, U+02BB-02BC, U+02C6, U+02DA,
U+02DC, U+0304, U+0308, U+0329, U+2000-206F, U+20AC, U+2122, U+2191,
U+2193, U+2212, U+2215, U+FEFF, U+FFFD;
}
/* Latin Extended: only when the page has Polish, Czech, Turkish, etc. */
@font-face {
font-family: "Source Serif 4";
src: url("/fonts/source-serif-4-latin-ext.woff2") format("woff2");
font-weight: 400;
font-display: swap;
unicode-range: U+0100-02BA, U+02BD-02C5, U+02C7-02CC, U+02CE-02D7, U+02DD-02FF,
U+1E00-1E9F, U+1EF2-1EFF, U+20A0-20AB, U+20AD-20C0, U+2C60-2C7F, U+A720-A7FF;
}
/* Cyrillic: only on Russian, Ukrainian, Bulgarian pages */
@font-face {
font-family: "Source Serif 4";
src: url("/fonts/source-serif-4-cyrillic.woff2") format("woff2");
font-weight: 400;
font-display: swap;
unicode-range: U+0301, U+0400-045F, U+0490-0491, U+04B0-04B1, U+2116;
}
An English page downloads only the first file. A Polish page with "ł" downloads the first two. This is how Google Fonts, Fontsource, and next/font deliver multi-script families efficiently, and it is particularly important for multilingual websites.
One caution: if a single character from a range appears on a page, the whole file for that range downloads. A stray Cyrillic letter in a user comment triggers the Cyrillic download. That is the correct behavior, but it explains surprise requests in your waterfall.
Subsetting in Frameworks and Services
- Next.js:
next/font/googleaccepts asubsetsarray, such assubsets: ["latin", "latin-ext"], and generatesunicode-rangerules. Only the listed subsets are preloaded. - Fontsource: packages ship with every subset as separate files and
unicode-rangerules; you can import a single subset, for example@fontsource/inter/latin-400.css. - Google Fonts API: serves per-subset files automatically. The
text=URL parameter returns a content-based subset for specific characters, which is useful for logos. - Adobe Fonts: offers dynamic subsetting in its web project settings, serving only the characters needed by each page.
Licensing and Subsetting
Subsetting creates a modified version of the font file, so check the license:
- SIL Open Font License fonts can be subset. Fonts with a Reserved Font Name technically require the modified version to use a different name, though web subsetting is widely practiced and the OFL FAQ discusses it specifically. When in doubt, use the subsets the designer or Google already distributes.
- Commercial licenses vary. Many web licenses explicitly allow subsetting for performance; some require the foundry's own tools or forbid modification. Read the EULA.
- Keep the
nametable intact, since it holds copyright and license information many licenses require you to preserve.
Avoiding Missing Glyphs
The main risk of subsetting is the tofu box or a jarring fallback glyph when a character is not in your subset. Prevent it:
- Use full script ranges for any text editors or users can change.
- Include typographic punctuation: curly quotes, en and em dashes, ellipsis, non-breaking space, and the bullet are all in the
U+2000-206Fblock. - Include currency symbols your business uses, such as €, £, ₹, or ৳.
- Keep a sensible fallback stack after the web font, so missing characters render in a similar system font rather than a random one.
- Test with a page containing every character you claim to support, including accented names like "Zoë", "Søren", and "Łukasz".
Font Subsetting FAQ
No. The glyphs you keep are identical to the originals. Subsetting only removes characters and data you do not use. Dropping hinting can slightly change rendering on low-resolution Windows screens, so test that option separately.
No. Subsetting removes content from the font, while compression formats such as WOFF2 encode the remaining content more efficiently. You should do both: subset first, then output WOFF2.
Only if you drop the layout features that contain them. Keep features such as kern, liga, and calt in the layout features option, or keep all features with an asterisk, and kerning and ligatures continue to work for the kept characters.
It tells the browser which characters a font file covers, so the browser downloads that file only when a page contains those characters. It lets you split one family into several script subsets that load on demand.
Yes. Google Fonts serves script-based subsets with unicode-range automatically, and next/font lets you choose which subsets to include and preload. You only need manual subsetting for fonts from other sources or for very specific content-based subsets.
It depends on the license. Many web font licenses allow subsetting for performance, but some forbid any modification or require the foundry's own tools. Read the license agreement before subsetting a commercial font.
Conclusion
Most font files are built to serve every possible language and use, which makes them far larger than any single website needs. Subsetting strips out the glyphs and features you do not use, often cutting file size by 60–90%, and that directly speeds up text rendering on slow connections.
Use script-based subsets for body text and content-based subsets only for fixed display text. Subset with pyftsubset (keeping the layout features your design depends on), output WOFF2, and split scripts into separate files with unicode-range so pages download only what they need. Check the license, test with accented and special characters, and you get faster fonts with no visible downside.
Here are some useful references for going deeper on font subsetting:
- fontTools Documentation: fontTools.subset — every
pyftsubsetoption, including layout feature and hinting controls. - MDN Web Docs: unicode-range — syntax and behavior of the descriptor.
- GitHub: glyphhanger by Zach Leatherman — crawling pages for used characters and subsetting fonts.
- web.dev: Best practices for fonts — subsetting and unicode-range in the context of Core Web Vitals.
- SIL: OFL FAQ — how the Open Font License treats modified and subset fonts.


