Skip to content
ASCII World

How AI Tokenization Works: From ASCII to Byte-Pair Encoding

How modern language models transform character sets and raw UTF-8 bytes into numerical subword vectors.

By the ASCII World team

From ASCII Encodings to Modern character sets

In 1963, the American Standards Association published ASA X3.4-1963, establishing 7-bit ASCII code points from 0 to 127 as the baseline standard for digital character processing. This system mapped every letter, number, and control symbol to a specific single-byte binary value, fitting neatly within standard 8-bit computational architectures. While the 7-bit ASCII standard served English text efficiently, expanding software globally required handling complex scripts and symbols beyond the initial 128 characters.

The introduction of Unicode in 1991 addressed this geographic constraint by assigning a unique Unicode code point to every character across human languages. To transmit these code points efficiently over legacy networks designed for ASCII characters, Ken Thompson and Rob Pike designed UTF-8 in 1992. UTF-8 uses a variable-length scheme where standard ASCII characters occupy a single byte identical to their original values, while non-ASCII characters occupy two, three, or four bytes. The variable-length encoding structure of UTF-8 provided backward compatibility with existing systems while expanding the potential character space to 1,114,112 code points.

The Core Problem: Converting Strings to Model Inputs

Neural networks operate on linear algebra principles, requiring numerical inputs in the form of vectors rather than raw string representations. Early natural language processing architectures experimented with two simplistic approaches to text input: word-level split tokenization and raw character-level processing. Both methods presented immediate computational bottlenecks when scaled to billion-parameter transformer models.

Word-level tokenization relies on splitting text by whitespace and punctuation. This method generates massive vocabulary tables containing millions of individual tokens to cover morphological variations like "run", "running", and "runner". Despite large vocabulary tables, word-level systems repeatedly fail when encountering unseen terms, misspellings, or domain-specific jargon, labeling them as unknown tokens. Conversely, character-level tokenization reduces vocabulary size down to the 256 possible byte values, eliminating unknown tokens entirely. However, character-level inputs lengthen sequence lengths by four to six times compared to word-level models. Because attention mechanism computational complexity scales quadratically with sequence length, processing long documents at the character level exhausts hardware context memory quickly, as explained in our guide on storing text in computer memory.

Byte-Pair Encoding: Merging Data Compression and Neural Inputs

In 1994, Philip Gage published a text compression algorithm named Byte-Pair Encoding (BPE) in the C/C++ Users Journal. Gage designed BPE to compress data by finding the most frequent adjacent pair of bytes in a document and replacing that pair with an unused single-byte value. The process repeated iteratively until no further frequent byte pairs remained or available unused byte values ran out.

Rico Sennrich, Barry Haddow, and Alexandra Birch adapted BPE for neural machine translation in 2015. Instead of compressing text to save disk space, they applied pair-merging logic to build subword vocabularies. Subword tokenization sits between character-level and word-level representations. Frequent words like "the" remain single tokens, while rare words like "unbelievable" break down into recognized subword components like "un", "believ", and "able".

In 2019, OpenAI modified subword BPE for GPT-2 by shifting the target unit from Unicode characters to raw UTF-8 bytes. Byte-Level BPE operates directly on raw byte streams rather than decoded text strings. Starting with a base vocabulary of 256 individual byte values, Byte-Level BPE ensures the model can tokenize any text string without encountering out-of-vocabulary errors, regardless of language, formatting, or special characters.

The Step-by-Step Mechanics of Tokenization

Building a Byte-Level BPE tokenizer requires a two-phase process: training the vocabulary on a target text corpus, and applying the learned merges to split raw text during inference.

During vocabulary training, the algorithm begins with a base vocabulary of 256 byte tokens, representing all possible single-byte values from 0x00 to 0xFF in hexadecimal representations. Next, the algorithm scans the corpus to record the frequency of every adjacent byte pair. The pair with the highest frequency merges into a new, single token and joins the vocabulary dictionary. The algorithm then re-scans the text, replacing all occurrences of that pair with the new token ID, and repeats this cycle until reaching a predefined target vocabulary size, such as 50,257 tokens for GPT-2 or 100,000 tokens for GPT-4's cl100k_base.

Iterative StepTarget PairFrequencyAction TakenNew Vocabulary Entry
Initial StateNoneN/AInitialize base byte vocabulary256 Byte Tokens (0-255)
Step 1't', 'h'1,420,500Merge adjacent bytes 't' and 'h'Token #256 ('th')
Step 2'th', 'e'980,200Merge token #256 with byte 'e'Token #257 ('the')
Step 3'i', 'n'850,100Merge adjacent bytes 'i' and 'n'Token #258 ('in')
Step 4' ', 'th'710,400Merge leading space byte with 'th'Token #259 (' th')

During inference, raw input strings are split into character or byte sequences. Regex rules often pre-segment input strings to prevent merges across distinct categories like punctuation and letters. For example, OpenAI tokenizers prevent space characters from merging across sentence boundaries or blending numbers with alphabetic characters. Once pre-segmented, the tokenizer applies trained merge rules in order of priority to generate the final array of token IDs sent to the model embedding layer.

Whitespace Handling and Language Token Density

Byte-level tokenizers treat space characters as distinct bytes rather than transparent separators. In GPT tokenizers, space characters prefixing words are included within subword tokens, represented visually by symbols like Ġ in HuggingFace tools or direct byte values in tiktoken. As a result, the string " token" and the standalone string "token" map to completely different token IDs.

This design choice creates significant token density variations across different languages. English ASCII text features highly compressed token ratios, averaging approximately 1.3 tokens per word because single-byte ASCII characters form frequent subword combinations during BPE training. Non-English languages using non-Latin scripts require multi-byte UTF-8 structures where a single character spans three or four bytes. If the BPE training corpus lacks sufficient non-English data, non-Latin words fail to merge into higher-order tokens, forcing the tokenizer to fall back on individual byte tokens.

A single word in Arabic, Chinese, or Hindi may require four to eight tokens, whereas the equivalent English word uses a single token. This disparity directly increases computational latency and API costs for non-English users, as large language models calculate processing limits and context window usage strictly based on total token count rather than character count.

References

  1. Gage, P. (1994). A New Algorithm for Data Compression. C/C++ Users Journal
  2. Sennrich, R., Haddow, B., & Birch, A. (2015). Neural Machine Translation of Rare Words with Subword Units. arXiv:1508.07909
  3. Radford, A., et al. (2019). Language Models are Unsupervised Multitask Learners (GPT-2)
  4. Unicode Consortium (2023). The Unicode Standard, Version 15.0