What is Unicode and why does text still break?
Unicode is a standard assigning a unique number — a code point — to every character across every writing system. It exists because the earlier situation was chaos: dozens of incompatible encodings, each covering a subset of characters, with no way to tell which a file used.
The crucial distinction that causes most confusion: Unicode is not an encoding. Unicode assigns numbers to characters. An encoding — UTF-8, UTF-16, UTF-32 — specifies how those numbers become bytes.
UTF-8 is the dominant encoding and deservedly so. It is variable width: ASCII characters take one byte, making it backward compatible with the entire existing world of ASCII text, while other characters take two to four bytes. It is self-synchronising and endianness-free.
Why text still breaks:
Mojibake — text decoded with the wrong encoding, producing ’ where an apostrophe should be. This happens when the encoding is not declared, declared wrongly, or assumed. The rule is that there is no such thing as plain text: bytes are meaningless without knowing the encoding.
A character is not a byte, and not a code point either. Three distinct concepts:
Code point — one Unicode number.
Code unit — one unit of the encoding, which may be a fraction of a character.
Grapheme cluster — what a user perceives as one character, which may be several code points. An emoji with a skin tone modifier, a family emoji, or an accented letter formed by combining marks are all single visible characters made of multiple code points.
This is why naive string length, reversal and truncation break — cutting a string at a byte or code point boundary can split a grapheme, producing garbage or a different emoji.
Normalisation. The same visible character can be represented differently — precomposed or as base plus combining mark — so two strings that look identical compare unequal. Normalise before comparing.
Case conversion is locale-dependent, and sorting genuinely differs by language.