Topic summary
Encoding

Extracted from the Wikipedia article Code.
Character encoding
A character encoding describes how character-based data (text) is encoded. Antiquated encoding systems used a fixed number of bits, ranging from 4 to 7, but modern systems use one or more 8-bitbytes for each character. ASCII, the dominate system for decades, uses one byte for each character, and therefore, can encode up to 256 different characters. To support natural languages with more characters, other systems were invented that use more than one byte or a variable number of bytes for each character. A writing system with a large character set such as Chinese, Japanese and Korean can be represented with a multibyte encoding. Early multibyte encodings were fixed-length, meaning that each character is represented by the same number of bytes, making them suitable for decoding via a lookup table. On the other hand, a variable-width encoding is more complex to decode since it cannot be decoded via a single lookup table and must be processed sequentially, but it supports a more efficient representation of a large character set by using a smaller representation for more commonly used characters. Today, UTF-8, an encoding of the Unicode character set, is the most common text encoding used on the Internet.