i

Where text breaks

49 classes · 217,414 characters

The Unicode line-breaking classes, counted by character.

The finding. Every design system assumes a word is a run of letters between spaces. That describes one character in eight. Ideographic characters — which break almost anywhere — outnumber alphabetic ones 6.4 to 1, and an entire class cannot be broken at all without a dictionary.

breaks anywhere 84.7%
breaks spaces 12.9%
breaks attaches 1.2%
breaks other 0.6%
breaks locale 0.4%
breaks dictionary 0.3%
What an engine must know

ClassNameCharactersWhat an engine must know
IDIdeographic 172,561 Breaks between almost any two characters. No spaces required, and none used.
XXUnknown 137,468 Unassigned codepoints, defaulting.
ALAlphabetic 26,954 Breaks at spaces — the Latin model, and the one every design system assumes.
H3Hangul LVT Syllable 10,773 Korean syllable block.
CMCombining Mark 2,512 Takes the class of the character it attaches to.
SGSurrogate 2,048 Not text on its own.
SAComplex Context Dependent 757 Cannot be broken without a dictionary or a trained model — Thai, Lao, Khmer, Myanmar.
AIAmbiguous 718 Resolved by locale, not by the character. The same codepoint breaks differently in two languages.
NUNumeric 705 Digits, which must not be split from their separators.
H2Hangul LV Syllable 399 Korean syllable block.
AKAK 329 other
BABreak After 263 A break opportunity follows.
ASAS 214 other
JTHangul T Jamo 137 Korean trailing consonant.
EBEB 134 other
JLHangul L Jamo 125 Korean leading consonant.
OPOP 95 other
JVHangul V Jamo 95 Korean vowel.
CLCL 94 other
HLHL 75 other
PRPR 67 other
CJConditional Japanese Starter 60 May or may not start a line depending on Japanese typographic preference.
BBBreak Before 55 A break opportunity precedes.
GLGL 41 other
EXEX 40 other
QUQU 39 other
POPO 38 other
NSNS 37 other
RIRI 26 other
HHHH 11 other
ISIS 10 other
VIVI 7 other
CPCP 6 other
ININ 6 other
APAP 6 other
EMEM 5 other
BKBK 4 other
B2B2 3 other
VFVF 2 other
WJWJ 2 other
LFLF 1 other
CRCR 1 other
SPSP 1 other
HYHY 1 other
SYSY 1 other
NLNL 1 other
ZWZW 1 other
ZWJZWJ 1 other
CBCB 1 other

Source. Unicode Character Database, LineBreak.txt (UAX #14) · Unicode Licence · retrieved 2026-07-31. This collection is reproduced.