Type something to search...
What Is Font Subsetting, and How Does It Reduce File Size?

What Is Font Subsetting, and How Does It Reduce File Size?

You download a font from a foundry for a client's English-language marketing site and drop the WOFF2 into the project. It weighs 310KB. The site's entire JavaScript bundle is smaller than that. Open the font in a glyph viewer and you find out why: it contains Latin, Latin Extended, Cyrillic, Greek, Vietnamese, hundreds of symbols, fractions, alternate numerals, and stylistic sets the design never uses. The site needs maybe 200 of the font's 2,500 glyphs.

Font subsetting is how you ship only the glyphs you need. This article explains what subsetting removes, how much it saves, how to subset with pyftsubset and glyphhanger, how unicode-range lets you split a font into on-demand pieces, and the mistakes that leave readers looking at empty boxes.

What Is Font Subsetting?

Font subsetting is the process of creating a smaller version of a font file that contains only a chosen set of characters (and the glyphs and features needed to render them). Everything else is stripped out.

A font file contains much more than glyph outlines:

Font componentWhat it holdsEffect of subsetting
Glyph outlines (glyf or CFF)The vector shapes of every letterRemoved for unused characters; the biggest saving
Character map (cmap)Mapping from Unicode code points to glyphsTrimmed to the kept characters
Layout features (GSUB, GPOS)Kerning, ligatures, alternates, small caps, mark positioningCan be kept, trimmed, or dropped
Hinting (fpgm, prep, cvt)Instructions for rendering at small sizes on low-res screensCan be dropped to save space
Variation data (gvar, fvar)Interpolation data for variable font axesKept, but scaled down with the outlines
Names and metadata (name)Font names, copyright, license textUsually kept; required by most licenses

The glyph outlines account for most of the file size, so removing scripts you do not use delivers most of the savings.

How Much Does It Save?

Savings depend on how many scripts the original font covers. Typical results for a well-equipped sans-serif in WOFF2:

VersionApproximate size
Full font, all scripts250–350KB
Latin plus Latin Extended80–120KB
Basic Latin subset (Google's "latin" range)30–60KB
Basic Latin, no hinting20–45KB
Uppercase A–Z only (for a logo or drop cap)5–10KB

A variable font with several axes stays larger, but the proportional savings are similar. Cutting 200KB from a font that renders your above-the-fold text can noticeably improve Largest Contentful Paint on mobile, and smaller files are far more likely to arrive inside the short window used by font-display: optional.

Subsetting by Script vs. by Content

There are two broad approaches:

  1. Script-based subsetting keeps entire Unicode ranges: all of Basic Latin, all of Latin Extended-A, all of Cyrillic. It is safe for user-generated and CMS content because any text in that script renders. This is what Google Fonts does.
  2. Content-based subsetting keeps only the exact characters found on your pages. It produces the smallest files but breaks the moment someone types a character you did not include. It suits fixed text: a logo wordmark, a navigation set in a display face, a single drop cap.

For body text on any site with editable content, use script-based subsetting. For decorative display faces used in a handful of fixed places, content-based subsetting is fine.

Subsetting with pyftsubset

pyftsubset is part of fontTools, the Python library maintained by the font engineering community and used by Google Fonts. Install it with Brotli support so it can write WOFF2:

pip install fonttools brotli

A Latin Subset for Body Text

pyftsubset SourceSerif4-Regular.ttf \
  --unicodes="U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD" \
  --layout-features="kern,liga,calt,ccmp,locl,mark,mkmk,onum,lnum,tnum,pnum,frac,sups,subs" \
  --flavor=woff2 \
  --output-file=source-serif-4-latin.woff2

The key options:

  • --unicodes lists the code points to keep. The range above is the same "latin" subset Google Fonts serves, covering English and most Western European languages, plus common punctuation, the euro sign, and the trademark symbol.
  • --layout-features lists the OpenType features to keep. Use "*" to keep them all. If you drop features like kern or liga, you lose kerning and ligatures in the subset. Keep any features your CSS turns on with font-feature-settings.
  • --flavor=woff2 writes compressed WOFF2 output.
  • --output-file sets the output name.

Other Useful Options

# Subset to the exact characters in a text file
pyftsubset Display.ttf --text-file=wordmark.txt --flavor=woff2 --output-file=display-wordmark.woff2

# Subset to specific characters inline
pyftsubset Display.ttf --text="ABCDEFGHIJKLMNOPQRSTUVWXYZ" --flavor=woff2 --output-file=display-caps.woff2

# Drop hinting for extra savings on modern high-DPI screens
pyftsubset Inter.ttf --unicodes-file=latin.txt --no-hinting --desubroutinize --flavor=woff2 --output-file=inter-latin.woff2

--no-hinting removes TrueType hinting instructions. Modern macOS, iOS, and Android ignore hinting, and Windows renders most fonts acceptably without it at typical body sizes on high-density displays. If many readers use Windows on low-resolution monitors, test before dropping hints.

--desubroutinize expands CFF subroutines in OpenType CFF fonts. It makes the raw font larger but often compresses better in WOFF2, so the final file is smaller.

Subsetting a Variable Font

pyftsubset handles variable fonts the same way; the variation tables are subset along with the outlines. If you only need part of an axis range, first restrict the axis with fontTools' instancer, then subset:

fonttools varLib.instancer "Inter[opsz,wght].ttf" wght=400:700 opsz=14:32 -o inter-400-700.ttf
pyftsubset inter-400-700.ttf --unicodes-file=latin.txt --layout-features="*" --flavor=woff2 --output-file=inter-latin-400-700.woff2

Limiting the weight axis from 100–900 to 400–700 can cut the variation data substantially. For more on the trade-offs, see our comparison of variable fonts vs. static fonts.

Finding the Characters You Use with glyphhanger

glyphhanger, a Node tool by Zach Leatherman, crawls pages and reports the Unicode code points they use. It can also run the subset for you, using fontTools under the hood.

npm install -g glyphhanger

# Report the characters used across a site
glyphhanger https://example.com --spider --spider-limit=50

# Subset a font to those characters and output WOFF2
glyphhanger https://example.com --spider --subset=Display.ttf --formats=woff2

Use it to audit what characters your content contains, then round the result up to full script ranges for body fonts. Its output is a great starting point for display fonts used only in fixed places.

Splitting Subsets with unicode-range

You do not have to choose one subset. The unicode-range descriptor lets you declare several files for the same family, and the browser downloads each file only if the page contains characters in its range.

/* Latin: downloaded on almost every page */
@font-face {
  font-family: "Source Serif 4";
  src: url("/fonts/source-serif-4-latin.woff2") format("woff2");
  font-weight: 400;
  font-display: swap;
  unicode-range: U+0000-00FF, U+0131, U+0152-0153, U+02BB-02BC, U+02C6, U+02DA,
    U+02DC, U+0304, U+0308, U+0329, U+2000-206F, U+20AC, U+2122, U+2191,
    U+2193, U+2212, U+2215, U+FEFF, U+FFFD;
}

/* Latin Extended: only when the page has Polish, Czech, Turkish, etc. */
@font-face {
  font-family: "Source Serif 4";
  src: url("/fonts/source-serif-4-latin-ext.woff2") format("woff2");
  font-weight: 400;
  font-display: swap;
  unicode-range: U+0100-02BA, U+02BD-02C5, U+02C7-02CC, U+02CE-02D7, U+02DD-02FF,
    U+1E00-1E9F, U+1EF2-1EFF, U+20A0-20AB, U+20AD-20C0, U+2C60-2C7F, U+A720-A7FF;
}

/* Cyrillic: only on Russian, Ukrainian, Bulgarian pages */
@font-face {
  font-family: "Source Serif 4";
  src: url("/fonts/source-serif-4-cyrillic.woff2") format("woff2");
  font-weight: 400;
  font-display: swap;
  unicode-range: U+0301, U+0400-045F, U+0490-0491, U+04B0-04B1, U+2116;
}

An English page downloads only the first file. A Polish page with "ł" downloads the first two. This is how Google Fonts, Fontsource, and next/font deliver multi-script families efficiently, and it is particularly important for multilingual websites.

One caution: if a single character from a range appears on a page, the whole file for that range downloads. A stray Cyrillic letter in a user comment triggers the Cyrillic download. That is the correct behavior, but it explains surprise requests in your waterfall.

Subsetting in Frameworks and Services

  • Next.js: next/font/google accepts a subsets array, such as subsets: ["latin", "latin-ext"], and generates unicode-range rules. Only the listed subsets are preloaded.
  • Fontsource: packages ship with every subset as separate files and unicode-range rules; you can import a single subset, for example @fontsource/inter/latin-400.css.
  • Google Fonts API: serves per-subset files automatically. The text= URL parameter returns a content-based subset for specific characters, which is useful for logos.
  • Adobe Fonts: offers dynamic subsetting in its web project settings, serving only the characters needed by each page.

Licensing and Subsetting

Subsetting creates a modified version of the font file, so check the license:

  • SIL Open Font License fonts can be subset. Fonts with a Reserved Font Name technically require the modified version to use a different name, though web subsetting is widely practiced and the OFL FAQ discusses it specifically. When in doubt, use the subsets the designer or Google already distributes.
  • Commercial licenses vary. Many web licenses explicitly allow subsetting for performance; some require the foundry's own tools or forbid modification. Read the EULA.
  • Keep the name table intact, since it holds copyright and license information many licenses require you to preserve.

Avoiding Missing Glyphs

The main risk of subsetting is the tofu box or a jarring fallback glyph when a character is not in your subset. Prevent it:

  1. Use full script ranges for any text editors or users can change.
  2. Include typographic punctuation: curly quotes, en and em dashes, ellipsis, non-breaking space, and the bullet are all in the U+2000-206F block.
  3. Include currency symbols your business uses, such as €, £, ₹, or ৳.
  4. Keep a sensible fallback stack after the web font, so missing characters render in a similar system font rather than a random one.
  5. Test with a page containing every character you claim to support, including accented names like "Zoë", "Søren", and "Łukasz".

Font Subsetting FAQ

No. The glyphs you keep are identical to the originals. Subsetting only removes characters and data you do not use. Dropping hinting can slightly change rendering on low-resolution Windows screens, so test that option separately.

No. Subsetting removes content from the font, while compression formats such as WOFF2 encode the remaining content more efficiently. You should do both: subset first, then output WOFF2.

Only if you drop the layout features that contain them. Keep features such as kern, liga, and calt in the layout features option, or keep all features with an asterisk, and kerning and ligatures continue to work for the kept characters.

It tells the browser which characters a font file covers, so the browser downloads that file only when a page contains those characters. It lets you split one family into several script subsets that load on demand.

Yes. Google Fonts serves script-based subsets with unicode-range automatically, and next/font lets you choose which subsets to include and preload. You only need manual subsetting for fonts from other sources or for very specific content-based subsets.

It depends on the license. Many web font licenses allow subsetting for performance, but some forbid any modification or require the foundry's own tools. Read the license agreement before subsetting a commercial font.

Conclusion

Most font files are built to serve every possible language and use, which makes them far larger than any single website needs. Subsetting strips out the glyphs and features you do not use, often cutting file size by 60–90%, and that directly speeds up text rendering on slow connections.

Use script-based subsets for body text and content-based subsets only for fixed display text. Subset with pyftsubset (keeping the layout features your design depends on), output WOFF2, and split scripts into separate files with unicode-range so pages download only what they need. Check the license, test with accented and special characters, and you get faster fonts with no visible downside.

Here are some useful references for going deeper on font subsetting:

  1. fontTools Documentation: fontTools.subset — every pyftsubset option, including layout feature and hinting controls.
  2. MDN Web Docs: unicode-range — syntax and behavior of the descriptor.
  3. GitHub: glyphhanger by Zach Leatherman — crawling pages for used characters and subsetting fonts.
  4. web.dev: Best practices for fonts — subsetting and unicode-range in the context of Core Web Vitals.
  5. SIL: OFL FAQ — how the Open Font License treats modified and subset fonts.
Share :

Related Posts

Ascenders, Descenders, and Baselines: The Anatomy of a Letterform

Ascenders, Descenders, and Baselines: The Anatomy of a Letterform

You align an icon next to a button label and it looks a pixel or two too high, no matter how you adjust vertical-align. You set overflow: hidden

Continue Reading
Are Google Fonts GDPR-Compliant?

Are Google Fonts GDPR-Compliant?

In late 2022, thousands of small business owners in Germany and Austria opened letters demanding a few hundred euros in "damages" because their websi

Continue Reading
How to Audit Web Font Performance with Lighthouse?

How to Audit Web Font Performance with Lighthouse?

A client sends you a screenshot of their PageSpeed Insights report: performance score 61, LCP 3.9 seconds, and a vague list of warnings. They want to

Continue Reading