Type something to search...
What Are Glyphs, and How Do They Differ from Characters?

What Are Glyphs, and How Do They Differ from Characters?

A product name field limits input to 20 characters. A customer in Bangladesh types their shop name in Bengali, and the form rejects it even though it looks short. Another customer pastes a name containing a family emoji, and the counter jumps by 11. Meanwhile, a designer asks why the font's specimen page says it has "over 2,000 glyphs" when the alphabet only has 26 letters. All of these confusions come from treating characters and glyphs as the same thing. They aren't, and once you understand the difference, a lot of font and text-handling behavior starts to make sense.

This article explains what characters and glyphs are, how a font turns one into the other, the many-to-one and one-to-many relationships between them, and why the distinction matters for CSS, font loading, fallback, and JavaScript string handling.

Characters vs. Glyphs: The Short Version

A character is an abstract unit of text with a meaning. It's what you store, search, and transmit. "Latin small letter a" is a character, identified in Unicode as code point U+0061.

A glyph is a specific visual shape a font uses to draw one or more characters. The "a" in Georgia, the "a" in Helvetica, and the single-storey "a" in a typical italic are three different glyphs for the same character.

AspectCharacterGlyph
What it isA unit of meaningA unit of visual form
Defined byUnicodeThe font designer
Identified byCode point, such as U+0061Glyph ID inside a specific font
Lives inYour HTML, database, and stringsThe font file
AffectsSearch, sorting, copy, screen readersAppearance only

The rule of thumb: text is made of characters; rendering is made of glyphs. Your markup should always contain the correct characters. The font decides which glyphs to draw.

What Unicode Defines

Unicode assigns a code point to every character in nearly every writing system, plus symbols and emoji. Code points range from U+0000 to U+10FFFF, organized into blocks: Basic Latin (U+0000–U+007F), Latin-1 Supplement (U+0080–U+00FF), Bengali (U+0980–U+09FF), and so on.

Unicode deliberately encodes characters, not appearances. That's why there's only one code point for "a" no matter how many designs exist, and why bold, italic, and font choice are left to styling. There are a few historical exceptions, like the "presentation forms" for ligatures such as U+FB01 (fi) and the "Mathematical Alphanumeric Symbols" block, but those exist for compatibility and specialized notation, not as a styling tool. Using them to fake bold or italic text in social posts breaks search and screen readers.

On the web, characters arrive as bytes in a specific encoding. Always declare UTF-8:

<meta charset="utf-8" />

What a Font Contains

A font file is essentially a collection of glyphs plus the tables that describe how to use them. In an OpenType font, the important ones are:

  • cmap (character map): maps Unicode code points to default glyph IDs. This is how the font says "for U+0061, draw glyph 68."
  • glyf or CFF: the actual outlines of each glyph.
  • GSUB (glyph substitution): rules that replace glyphs with other glyphs, such as ligatures, small caps, alternates, and contextual forms.
  • GPOS (glyph positioning): rules for kerning and mark placement.
  • hmtx: the advance width of each glyph.

That's why a font for Latin text can contain thousands of glyphs. Beyond the basic alphabet, a well-equipped family includes accented letters, small caps, old-style and lining figures, tabular figures, fractions, ligatures, stylistic alternates, arrows, currency symbols, and glyphs for other scripts.

You can inspect these tables directly with fontTools:

pip install fonttools brotli

# Count glyphs and mapped characters
python3 -c "
from fontTools.ttLib import TTFont
f = TTFont('Inter-Regular.woff2')
print('glyphs:', len(f.getGlyphOrder()))
print('mapped code points:', len(f.getBestCmap()))
"

# Dump the cmap table to readable XML
ttx -t cmap -o cmap.xml Inter-Regular.ttf

The number of glyphs is almost always larger than the number of mapped code points, because many glyphs (alternates, ligatures, small caps) are only reachable through substitution rules, not directly from a character.

How Text Becomes Glyphs: Shaping

Between your HTML and the pixels on screen, the browser runs a text shaping engine (HarfBuzz in Chrome, Firefox, and Android; Core Text in Safari). Shaping takes a run of characters and a font, and produces a sequence of positioned glyphs. In simplified form:

  1. Look up each character's default glyph in the cmap.
  2. Apply GSUB substitutions the active features ask for: ligatures, contextual forms, any features you enabled in CSS.
  3. Apply GPOS positioning: kerning, mark attachment for accents and diacritics.
  4. Output glyph IDs with x and y positions.

For Latin text, this mostly looks like a one-to-one mapping. For scripts like Arabic, Devanagari, or Bengali, shaping is essential: the same character takes different forms depending on its neighbors, and clusters of characters combine into single visual units.

The Relationships Between Characters and Glyphs

One Character, One Glyph

The common case. U+0041 "A" maps to the font's A glyph.

Many Characters, One Glyph

Ligatures combine several characters into one glyph: "f" + "i" becomes a single fi glyph. Bengali and Devanagari conjuncts combine consonants and a virama into one shape. The post on ligatures covers the Latin side of this.

One Character, Several Glyphs

A single character can have several possible glyphs in the same font:

  • Positional forms: Arabic letters have isolated, initial, medial, and final forms.
  • Stylistic alternates: a font may offer a single-storey "a" or a straight-legged "R" via stylistic sets.
  • Small caps: a lowercase letter shown as a small capital.
  • Figure styles: the digit 1 as a lining, old-style, tabular, or proportional glyph.

You choose between these with CSS, and the underlying character never changes:

/* Small caps glyphs, still lowercase characters in the DOM */
.label {
  font-variant-caps: all-small-caps;
}

/* Tabular figures so digits line up in columns */
.price-table td {
  font-variant-numeric: tabular-nums lining-nums;
}

/* Stylistic set 01, if the font defines one */
.brand {
  font-feature-settings: "ss01" 1;
}

Number styles get their own post on tabular figures for numbers and tables.

One Character, Zero Glyphs

Some characters have no visible glyph: the zero-width joiner (U+200D), the soft hyphen (U+00AD) when not at a line break, and various control characters.

Several Characters, One Visual Unit

Accented letters can be stored two ways. "é" can be the single precomposed character U+00E9, or "e" (U+0065) followed by the combining acute accent (U+0301). Both look identical when rendered, but they're different character sequences. The font draws either a precomposed glyph or a base glyph with a positioned mark.

Why the Difference Matters for Developers

String Length Is Not What Users See

JavaScript strings are sequences of UTF-16 code units, not characters and certainly not glyphs. Emoji, many CJK characters, and combining sequences break naive length checks:

const name = "👨‍👩‍👧‍👦";
console.log(name.length); // 11 UTF-16 code units
console.log([...name].length); // 7 code points

const segmenter = new Intl.Segmenter("en", { granularity: "grapheme" });
console.log([...segmenter.segment(name)].length); // 1 grapheme cluster

A grapheme cluster is what a user perceives as a single character. If you validate or truncate user input, count graphemes with Intl.Segmenter, which is supported in all modern browsers and Node.js.

function truncateGraphemes(text: string, max: number, locale = "en"): string {
  const segmenter = new Intl.Segmenter(locale, { granularity: "grapheme" });
  const parts = Array.from(segmenter.segment(text), (s) => s.segment);
  return parts.length > max ? parts.slice(0, max).join("") + "…" : text;
}

Normalize Before Comparing

Because "é" can be stored two ways, string comparisons can fail on text that looks identical. Normalize to NFC before storing or comparing:

const a = "café";
const b = "café";
console.log(a === b); // false
console.log(a.normalize("NFC") === b.normalize("NFC")); // true

Missing Glyphs and Fallback

If a font has no glyph for a character, the browser doesn't give up. It moves down the font-family stack and then to system fonts until it finds one that covers the character. If nothing does, you see the font's .notdef glyph, usually an empty box nicknamed "tofu."

Fallback happens character by character, which is why a heading in a custom Latin font can suddenly switch to a system font for a single "₹" or "→" that the custom font lacks. It looks broken because two different designs are mixed in one word. Check coverage for every symbol your content uses, not just the alphabet:

python3 -c "
from fontTools.ttLib import TTFont
cmap = TTFont('Inter-Regular.woff2').getBestCmap()
for ch in '₹€£→✓':
    print(ch, hex(ord(ch)), 'yes' if ord(ch) in cmap else 'MISSING')
"

unicode-range Works on Characters

The @font-face unicode-range descriptor tells the browser which code points a font file covers, so it only downloads the file when the page actually uses those characters. This is how Google Fonts splits families into per-script files:

@font-face {
  font-family: "Noto Sans Bengali";
  src: url("/fonts/noto-sans-bengali.woff2") format("woff2");
  unicode-range: U+0951-0952, U+0964-0965, U+0980-09FE, U+200C-200D;
  font-display: swap;
}

unicode-range is about characters, not glyphs. A file can contain far more glyphs than the code points listed, because shaping produces them from those characters. When you subset a font, you choose characters, and the subsetter keeps the glyphs those characters can reach through GSUB. See font subsetting for the full workflow.

Screen Readers, Search, and Copy Use Characters

Assistive technology, browser find, search engines, and the clipboard all work with characters. That's why you should never use glyph tricks to change meaning: don't use Unicode "bold" math letters for emphasis, and don't use icon fonts that map pictures to letter code points without accessible labels.


Glyphs and Characters FAQ

Not exactly. A letter is a character with a meaning, like lowercase a. A glyph is one particular drawing of it in one particular font. The same letter can be drawn by many different glyphs, and one glyph can represent several letters, as with ligatures.

Fonts include alternate drawings that are reached through substitution rules rather than direct character mappings. Small caps, ligatures, stylistic alternates, figure styles, and positional forms all add glyphs without adding new characters.

That is the notdef glyph, often called tofu. It appears when neither your chosen fonts nor any system font contains a glyph for the character. Adding a font that covers that script or symbol, such as a Noto family, fixes it.

No. Properties like font-variant-caps and font-feature-settings only change which glyphs are drawn. The characters in the DOM stay the same, so copy and paste, search, and screen readers behave as if the styling were not there.

Count grapheme clusters, which match what users perceive as single characters. In JavaScript, use Intl.Segmenter with grapheme granularity rather than string length, which counts UTF-16 code units and overcounts emoji and many scripts.

A glyph ID is the index of a glyph inside a specific font file. It has no meaning outside that font, which is why the same character can map to glyph 68 in one font and glyph 412 in another. Text is never stored as glyph IDs.

Conclusion

Characters are the units of meaning that Unicode defines and that your HTML, database, and JavaScript store. Glyphs are the shapes a font uses to draw them. A shaping engine sits between the two, mapping characters to glyphs through the font's cmap, then applying substitutions and positioning. The relationship can be one-to-one, many-to-one, one-to-many, or something more complex, depending on the script and the features you enable.

For day-to-day work, the practical lessons are clear. Store correct characters and let CSS choose glyphs. Check glyph coverage for every symbol your content uses. Use unicode-range and subsetting based on characters. Count graphemes, not string length, when validating input, and normalize text before comparing it.

Here are some useful references for going deeper on glyphs and characters:

  1. Unicode Consortium: Unicode Standard Annex #29: Text Segmentation — the official definition of grapheme clusters.
  2. MDN Web Docs: Intl.Segmenter — counting and splitting text by grapheme, word, or sentence.
  3. MDN Web Docs: unicode-range — how browsers use code point ranges to choose font files.
  4. Microsoft Typography: OpenType cmap table — how fonts map characters to glyphs.
  5. W3C: Character Model for the World Wide Web: String Matching — guidance on Unicode normalization and string comparison.
Share :

Related Posts

Ascenders, Descenders, and Baselines: The Anatomy of a Letterform

Ascenders, Descenders, and Baselines: The Anatomy of a Letterform

You align an icon next to a button label and it looks a pixel or two too high, no matter how you adjust vertical-align. You set overflow: hidden

Continue Reading
Are Google Fonts GDPR-Compliant?

Are Google Fonts GDPR-Compliant?

In late 2022, thousands of small business owners in Germany and Austria opened letters demanding a few hundred euros in "damages" because their websi

Continue Reading
How to Audit Web Font Performance with Lighthouse?

How to Audit Web Font Performance with Lighthouse?

A client sends you a screenshot of their PageSpeed Insights report: performance score 61, LCP 3.9 seconds, and a vague list of warnings. They want to

Continue Reading