Type something to search...
How to Display Bengali, Hindi, and Other Non-Latin Scripts on the Web?

How to Display Bengali, Hindi, and Other Non-Latin Scripts on the Web?

A news publisher in Dhaka moves its archive to a new website, and half the Bengali articles render as gibberish like "Avgvi †mvbvi evsjv". The other half display real Bengali letters, but vowel signs float away from their consonants and little dotted circles appear in the middle of words. A Hindi site built by the same agency has a different problem: the text is correct, but the self-hosted font was subset for performance, and now conjuncts like क्ष and श्र break into separate pieces with visible halant marks. None of these are browser bugs. They come from legacy encodings, missing fonts, and tooling that treats complex scripts like Latin.

This article explains what makes scripts like Bengali and Devanagari different, how to make sure your text is real Unicode, which fonts to use, how to load and subset them without breaking shaping, and the CSS adjustments these scripts need. The same principles apply to Tamil, Telugu, Gujarati, Thai, Khmer, Myanmar, and other complex scripts.

What Makes These Scripts "Complex"

Bengali (Bangla), Devanagari (used for Hindi, Marathi, Nepali, and Sanskrit), and their relatives are abugidas: each consonant carries an inherent vowel, and other vowels are written as marks attached to the consonant. Rendering them requires text shaping, where the font and the browser's shaping engine transform a sequence of Unicode characters into the correct arrangement of glyphs.

Several things happen during shaping that do not happen in Latin text:

  • Conjuncts. A consonant, the virama (hasanta in Bengali, halant in Hindi), and another consonant combine into a single ligature. Bengali ক + ্ + ষ becomes ক্ষ. Hindi क + ् + ष becomes क्ष.
  • Reordering. Some vowel signs are typed after the consonant but displayed before it. In কি and कि, the vowel sign is stored second but drawn on the left.
  • Reph. A ra at the start of a cluster becomes a small mark above the following consonant, as in র্ক and र्क.
  • Positioning. Marks above and below the base letter must be placed precisely, often depending on the consonant's shape.
  • The headline. Devanagari and Bengali letters hang from a horizontal stroke (shirorekha in Devanagari, matra in Bengali) that should connect across letters in a word.

All of this logic lives in the font's OpenType layout tables (GSUB for substitutions and GPOS for positioning) and in the browser's shaper. Chrome and Firefox use HarfBuzz, and Safari uses Core Text; all handle Indic scripts well when the font is correct. Your job is to give them proper Unicode text and a proper font, and then not break either.

ScriptUnicode blockLanguagesNotable features
BengaliU+0980–09FFBengali, Assamese, ManipuriConjuncts, reph, partial headline
DevanagariU+0900–097FHindi, Marathi, Nepali, SanskritConjuncts, reph, continuous headline
GujaratiU+0A80–0AFFGujaratiNo headline
TamilU+0B80–0BFFTamilFewer conjuncts, long words
ThaiU+0E00–0E7FThaiNo spaces between words
MyanmarU+1000–109FBurmeseStacked consonants, no word spaces

Step 1: Make Sure the Text Is Real Unicode

The "Avgvi †mvbvi" problem is the most common Bengali web issue, and no CSS will fix it. For decades, Bengali in Bangladesh was typed with the Bijoy keyboard and ANSI fonts such as SutonnyMJ. Those fonts draw Bengali shapes on top of Latin code points, so the stored text is actually Latin characters. With the original font installed, it looks like Bengali. Without it, or with any normal web font, readers see the underlying Latin letters. Hindi has the same legacy with fonts like Kruti Dev, and Burmese with Zawgyi.

Check your content:

// Run in the browser console on a page with suspect text
const text = document.querySelector("article").textContent;
const bengali = /[ঀ-৿]/.test(text);
const devanagari = /[ऀ-ॿ]/.test(text);
console.log({ bengali, devanagari });

If text that looks like Bengali on an editor's machine returns false, it is legacy-encoded. Convert it to Unicode before publishing, using an established Bijoy-to-Unicode or Kruti Dev-to-Unicode converter, and batch-convert archives in the database rather than patching them with legacy fonts on the front end. Embedding SutonnyMJ as a web font may make text look right, but it stays unsearchable, untranslatable, unreadable by screen readers, and broken for copy and paste.

Then make sure the page declares UTF-8 and the database stores it:

<meta charset="utf-8" />

For MySQL-backed systems like WordPress, the tables should use utf8mb4.

Step 2: Set the lang Attribute

Set lang on the root element and on any block in a different language:

<html lang="bn">
  ...
  <p lang="hi">यह वाक्य हिंदी में है।</p>
  <p lang="mr">हे वाक्य मराठीत आहे.</p>
</html>

lang matters for more than accessibility here. Marathi and Hindi both use Devanagari, but Marathi prefers different shapes for some letters, such as ल and श. Fonts that support both deliver the right shapes through the OpenType locl feature, which the browser only applies when lang="mr" is set. Screen readers also need lang to pick the correct voice. Common codes: bn (Bengali), as (Assamese), hi (Hindi), mr (Marathi), ne (Nepali), ta (Tamil), te (Telugu), gu (Gujarati), th (Thai), my (Burmese).

Step 3: Choose Fonts Built for the Script

Do not assume a popular Latin font covers Indic scripts. Choose families designed for the script by experienced type designers, with complete conjunct sets. Good options available on Google Fonts:

Bengali

  • Noto Sans Bengali and Noto Serif Bengali: broad coverage, reliable shaping, many weights
  • Hind Siliguri: clean UI sans with Latin companion
  • Anek Bangla: variable width and weight
  • Tiro Bangla: traditional text face for long reading
  • Baloo Da 2: rounded display style

Devanagari

  • Noto Sans Devanagari and Noto Serif Devanagari
  • Mukta and Hind: popular UI sans faces with Latin
  • Poppins: includes Devanagari alongside its geometric Latin
  • Tiro Devanagari Hindi and Tiro Devanagari Marathi: language-specific text faces
  • Anek Devanagari: variable width and weight

Many Indic families include a matched Latin set. That matters, because Bengali and Hindi pages almost always contain English words, brand names, and numbers. If you use a separate Latin font, pair them by stroke weight and apparent size; the general approach is in choosing fonts for multilingual websites.

Test candidates with strings that exercise shaping, not just simple words:

Bengali:  ক্ষ্ম  স্ত্র  র্ক  কি  কৌ  ঙ্ক্ষ  শ্রী  হৃ  জ্ঞ
Hindi:    क्ष  त्र  ज्ञ  श्र  द्ध  र्क  कि  कृ  ह्म  ट्ठ
Marathi:  ल  श  झ  (compare with lang="hi")

If a test string shows a dotted circle (◌) next to a mark, the font is missing a glyph or a shaping rule for that sequence, or the sequence itself is invalid Unicode. Either way, it is a red flag for that font.

Step 4: Load the Fonts Without Breaking Shaping

With next/font

next/font/google handles subsetting correctly for Google-hosted families. Request the script subset plus Latin:

// src/app/fonts.ts
import { Noto_Sans_Bengali, Hind } from "next/font/google";

export const bengali = Noto_Sans_Bengali({
  subsets: ["bengali", "latin"],
  weight: ["400", "600", "700"],
  variable: "--font-bengali",
  display: "swap",
});

export const hindi = Hind({
  subsets: ["devanagari", "latin"],
  weight: ["400", "600"],
  variable: "--font-hindi",
  display: "swap",
});

Apply the variables on html and reference them in CSS or a Tailwind theme. The full workflow is in how to use next/font in Next.js.

When Self-Hosting

Subsetting is where most self-hosted Indic fonts get broken. Two rules:

  1. Subset by Unicode range, never by sample text. Tools that subset to "the characters in this paragraph" drop conjunct glyphs that are not directly mapped to any character, because those glyphs are only reachable through GSUB substitutions.
  2. Keep all layout features. Dropping features like akhn, blwf, half, pres, abvs, and blws breaks conjuncts and mark positioning.

With pyftsubset, both are handled when you pass full ranges and keep every feature:

pip install fonttools brotli

# Bengali subset
pyftsubset NotoSansBengali-Regular.ttf \
  --unicodes="U+0980-09FF,U+0951-0952,U+0964-0965,U+200C-200D,U+20B9,U+25CC" \
  --layout-features='*' \
  --flavor=woff2 \
  --output-file=noto-sans-bengali-400.woff2

# Devanagari subset
pyftsubset NotoSansDevanagari-Regular.ttf \
  --unicodes="U+0900-097F,U+1CD0-1CF9,U+200C-200D,U+20A8,U+20B9,U+25CC,U+A830-A839,U+A8E0-A8FF" \
  --layout-features='*' \
  --flavor=woff2 \
  --output-file=noto-sans-devanagari-400.woff2

The extra code points matter: U+0964 and U+0965 are the danda and double danda (sentence-ending punctuation shared by Bengali and Devanagari), U+200C and U+200D are the zero-width non-joiner and joiner that control conjunct formation, U+20B9 is the rupee sign, and U+25CC is the dotted circle that shaping engines insert for invalid sequences. pyftsubset follows GSUB substitutions from the characters you keep, so conjunct glyphs are retained automatically. The general technique is covered in what font subsetting is.

Register the faces with matching unicode-range values so pages without Bengali or Devanagari never download them:

@font-face {
  font-family: "Site Bengali";
  src: url("/fonts/noto-sans-bengali-400.woff2") format("woff2");
  font-weight: 400;
  font-display: swap;
  unicode-range: U+0980-09FF, U+0951-0952, U+0964-0965, U+200C-200D, U+20B9, U+25CC;
}

Expect Indic font files to be larger than Latin ones, often 60–150KB per weight in WOFF2, because of the many conjunct glyphs. Load two or three weights at most, or use a variable version where one exists.

Step 5: Adjust Size and Line Height

At the same nominal font-size, Bengali and Devanagari look smaller than Latin text, because their main letter bodies occupy less of the em square, while marks above and below need extra room. Practical defaults:

:lang(bn) {
  font-family: "Site Bengali", "Noto Sans Bengali", system-ui, sans-serif;
  font-size: 1.125rem;
  line-height: 1.8;
}

:lang(hi),
:lang(mr),
:lang(ne) {
  font-family: var(--font-hindi), "Noto Sans Devanagari", system-ui, sans-serif;
  font-size: 1.0625rem;
  line-height: 1.75;
}

:lang(bn) :is(h1, h2, h3),
:lang(hi) :is(h1, h2, h3) {
  line-height: 1.35;
}

Tight Latin-style heading leading, like 1.1, makes vowel signs above one line collide with marks below the previous one. Keep body text at 1.7 or more and headings around 1.3 to 1.4. These values also comfortably exceed the WCAG 1.4.12 baseline of 1.5. Compare the rendered size to English text on the same page and adjust until they look balanced rather than matching the numbers.

Step 6: Avoid Latin-Only CSS Habits

Several common styles damage Indic text:

CSS habitEffect on Bengali / DevanagariDo instead
letter-spacing on headingsBreaks the connected headline into separate segmentsletter-spacing: 0
font-style: italicSynthetic slant; no true italics in most Indic fontsUse weight or color for emphasis
Missing bold weightBrowser synthesizes smeared boldLoad real bold or font-synthesis: none
text-transform: uppercaseNo effect on Indic letters, only on embedded LatinDo not rely on it for hierarchy
text-align: justifyUneven word gaps in long wordstext-align: start
line-height: 1.2 for body textMarks collide between lines1.7 or more
:lang(bn),
:lang(hi),
:lang(mr) {
  letter-spacing: 0;
  font-synthesis: none;
}

:lang(bn) :is(em, i),
:lang(hi) :is(em, i) {
  font-style: normal;
  font-weight: 600;
}

Truncation and Character Counting

JavaScript string methods count UTF-16 code units, not visible characters. Cutting a Bengali or Hindi string with slice can split a conjunct and leave a dangling virama or a dotted circle. Use Intl.Segmenter to cut at grapheme boundaries:

function truncate(text: string, maxGraphemes: number, locale = "bn"): string {
  const segmenter = new Intl.Segmenter(locale, { granularity: "grapheme" });
  const graphemes = Array.from(segmenter.segment(text), (s) => s.segment);
  return graphemes.length <= maxGraphemes
    ? text
    : graphemes.slice(0, maxGraphemes).join("") + "…";
}

For multi-line truncation in the browser, CSS line-clamp is safer, because the browser truncates at line boxes rather than inside clusters.

Numerals

Bengali has its own digits (০১২৩৪৫৬৭৮৯) and Devanagari has its own (०१२३४५६७८९). Which ones appear depends on your content and formatting code. Intl.NumberFormat follows locale conventions:

new Intl.NumberFormat("bn-BD").format(1234567); // "১২,৩৪,৫৬৭"
new Intl.NumberFormat("hi-IN").format(1234567); // "12,34,567"
new Intl.NumberFormat("hi-IN-u-nu-deva").format(1234567); // "१२,३४,५६७"

Note the Indian grouping (12,34,567) in both. Make sure your font includes the digits you output.

Other Scripts: What Changes

The same principles apply across complex scripts, with a few script-specific rules:

  • Thai, Lao, Khmer, and Myanmar do not use spaces between words. Browsers use dictionaries to find line-break opportunities, so do not insert manual breaks, and set lang so the right dictionary is used. These scripts also need generous line height for stacked marks.
  • Burmese has its own legacy-encoding problem with Zawgyi; convert content to Unicode as with Bijoy.
  • Tamil and Malayalam have long words that can overflow narrow columns; use overflow-wrap: break-word on containers.
  • Arabic, Hebrew, Persian, and Urdu add right-to-left layout on top of shaping, covered in handling right-to-left typography in CSS.
  • Chinese, Japanese, and Korean have huge character sets; Korean text usually needs word-break: keep-all to avoid breaking inside words.

Testing Checklist

  1. Confirm the content is Unicode (the console check above returns true).
  2. Confirm lang is set on html and on embedded passages.
  3. Render the shaping test strings in Chrome, Firefox, and Safari, on desktop and mobile.
  4. Look for dotted circles, detached vowel signs, or visible viramas where conjuncts should form.
  5. Check headings for collisions between lines and broken headlines.
  6. Check bold and emphasis for synthetic styles.
  7. Test truncated excerpts and search snippets for split clusters.
  8. Copy text from the page and paste it into a plain text editor to confirm it is real Bengali or Devanagari.

Non-Latin Scripts FAQ

The text was typed with a legacy ANSI system such as Bijoy with SutonnyMJ, which draws Bengali shapes over Latin code points. Without that exact font, the underlying Latin letters appear. Convert the content to Unicode Bengali with a Bijoy-to-Unicode converter.

The shaping engine inserts a dotted circle when a vowel sign or mark has no valid base to attach to. It means the character sequence is invalid, or the font is missing the glyph or rule needed to shape it.

Noto Sans Bengali and Noto Serif Bengali are reliable defaults with broad coverage and correct shaping. Hind Siliguri works well for interfaces, Tiro Bangla for long reading, and Anek Bangla when you want a variable font.

The subsetting tool probably dropped glyphs that are only reachable through OpenType substitutions, or removed layout features. Subset by full Unicode ranges, keep all layout features, and include the zero-width joiner and non-joiner.

Usually yes. At the same font size, Bengali and Devanagari look smaller than Latin text, so many sites set them five to fifteen percent larger with a body line height of about 1.7 to 1.9.

Avoid it. Letter spacing breaks the continuous headline stroke that connects Devanagari letters in a word, which looks broken to native readers. Use weight and size for emphasis instead.

Conclusion

Displaying Bengali, Hindi, and other complex scripts correctly depends on three things working together: real Unicode text, fonts designed for the script with complete shaping tables, and a browser that is allowed to do its shaping job. Most failures come from legacy encodings like Bijoy and Kruti Dev, from fonts that lack conjuncts, or from subsetting and string handling that treat these scripts like Latin.

Fix the content first, set lang everywhere, choose a well-made family like Noto, Hind, Mukta, or Tiro, and subset only by full Unicode ranges with all layout features kept. Then adjust the CSS for the script: larger sizes, line heights around 1.7 to 1.9, no letter spacing, no synthetic italics, and grapheme-aware truncation. Test with real shaping strings across browsers, and your pages will read as naturally to a reader in Dhaka or Delhi as they do in English.

Here are some useful references for going deeper on complex scripts on the web:

  1. W3C Internationalization: Indic Layout Requirements — the W3C note on line breaking, shaping, and typography for Indian scripts.
  2. Microsoft Typography: Developing OpenType fonts for Bengali script — how shaping, conjuncts, and features work in Bengali.
  3. Microsoft Typography: Developing OpenType fonts for Devanagari script — the equivalent reference for Hindi and other Devanagari languages.
  4. MDN Web Docs: Intl.Segmenter — grapheme-aware text segmentation for truncation and counting.
  5. Google Fonts: Noto — Noto families for Bengali, Devanagari, and nearly every other script.
Share :

Related Posts

Ascenders, Descenders, and Baselines: The Anatomy of a Letterform

Ascenders, Descenders, and Baselines: The Anatomy of a Letterform

You align an icon next to a button label and it looks a pixel or two too high, no matter how you adjust vertical-align. You set overflow: hidden

Continue Reading
Are Google Fonts GDPR-Compliant?

Are Google Fonts GDPR-Compliant?

In late 2022, thousands of small business owners in Germany and Austria opened letters demanding a few hundred euros in "damages" because their websi

Continue Reading
How to Audit Web Font Performance with Lighthouse?

How to Audit Web Font Performance with Lighthouse?

A client sends you a screenshot of their PageSpeed Insights report: performance score 61, LCP 3.9 seconds, and a vague list of warnings. They want to

Continue Reading