Where text breaks
49 classes · 217,414 charactersThe Unicode line-breaking classes, counted by character.
The finding. Every design system assumes a word is a run of letters between spaces. That describes one character in eight. Ideographic characters — which break almost anywhere — outnumber alphabetic ones 6.4 to 1, and an entire class cannot be broken at all without a dictionary.
| Class | Name | Characters | What an engine must know |
|---|---|---|---|
| ID | Ideographic | 172,561 | Breaks between almost any two characters. No spaces required, and none used. |
| XX | Unknown | 137,468 | Unassigned codepoints, defaulting. |
| AL | Alphabetic | 26,954 | Breaks at spaces — the Latin model, and the one every design system assumes. |
| H3 | Hangul LVT Syllable | 10,773 | Korean syllable block. |
| CM | Combining Mark | 2,512 | Takes the class of the character it attaches to. |
| SG | Surrogate | 2,048 | Not text on its own. |
| SA | Complex Context Dependent | 757 | Cannot be broken without a dictionary or a trained model — Thai, Lao, Khmer, Myanmar. |
| AI | Ambiguous | 718 | Resolved by locale, not by the character. The same codepoint breaks differently in two languages. |
| NU | Numeric | 705 | Digits, which must not be split from their separators. |
| H2 | Hangul LV Syllable | 399 | Korean syllable block. |
| AK | AK | 329 | other |
| BA | Break After | 263 | A break opportunity follows. |
| AS | AS | 214 | other |
| JT | Hangul T Jamo | 137 | Korean trailing consonant. |
| EB | EB | 134 | other |
| JL | Hangul L Jamo | 125 | Korean leading consonant. |
| OP | OP | 95 | other |
| JV | Hangul V Jamo | 95 | Korean vowel. |
| CL | CL | 94 | other |
| HL | HL | 75 | other |
| PR | PR | 67 | other |
| CJ | Conditional Japanese Starter | 60 | May or may not start a line depending on Japanese typographic preference. |
| BB | Break Before | 55 | A break opportunity precedes. |
| GL | GL | 41 | other |
| EX | EX | 40 | other |
| QU | QU | 39 | other |
| PO | PO | 38 | other |
| NS | NS | 37 | other |
| RI | RI | 26 | other |
| HH | HH | 11 | other |
| IS | IS | 10 | other |
| VI | VI | 7 | other |
| CP | CP | 6 | other |
| IN | IN | 6 | other |
| AP | AP | 6 | other |
| EM | EM | 5 | other |
| BK | BK | 4 | other |
| B2 | B2 | 3 | other |
| VF | VF | 2 | other |
| WJ | WJ | 2 | other |
| LF | LF | 1 | other |
| CR | CR | 1 | other |
| SP | SP | 1 | other |
| HY | HY | 1 | other |
| SY | SY | 1 | other |
| NL | NL | 1 | other |
| ZW | ZW | 1 | other |
| ZWJ | ZWJ | 1 | other |
| CB | CB | 1 | other |
Source. Unicode Character Database, LineBreak.txt (UAX #14) · Unicode Licence · retrieved 2026-07-31. This collection is reproduced.