Learning path

Full curriculum

Full curriculum

Unit content

Text encodings

Computers store text as bytes, but a byte sequence only becomes text when an encoding defines how those bytes map to characters.

Characters and code points

Unicode assigns abstract code points to characters from writing systems around the world. A code point is commonly written in a form such as

$$\text{U+0041}$$

for the character A.

A code point is not necessarily one byte. An encoding must still specify how the code point is represented as bytes.

ASCII

ASCII is an older encoding for a small set of English letters, digits, punctuation and control characters. Its values $0$ through $127$ are also the first part of Unicode.

UTF-8

UTF-8 encodes Unicode code points using one to four bytes. ASCII characters retain their one-byte representations, while other characters use longer sequences.

This variable-length design makes UTF-8 compact for ASCII-heavy text while supporting the full Unicode repertoire.

Encoding mismatches

If bytes written under one encoding are decoded under another, the displayed text can become corrupted or nonsensical. The bytes themselves have not changed; their interpretation has.

Text handling therefore always involves both the sequence of characters we mean and the byte encoding used to represent them.