AI & MLDeep
Intermediate

Tokenization in NLP

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Static Playbook
Static guide & reference tables

The first step nobody thinks about

Before a language model can do anything with text, it has to convert that text into numbers. Tokenization is the process of breaking raw text into smaller units — tokens — and mapping each one to an integer ID that the model's embedding layer understands.

It sounds like a boring preprocessing detail, but tokenization quietly shapes almost everything about how an LLM behaves: how much you pay per API call, how much text fits in a context window, why the model is bad at counting letters in a word, and why it sometimes stumbles on certain languages or code.

Why not just split on words or characters?

The naive options both fail. Word-level tokenization (splitting on spaces) creates a vocabulary problem: English alone has hundreds of thousands of word forms, and the model would have no way to handle a word it never saw during training — a rare name, a typo, a new slang term all become an unknown token.

Character-level tokenization solves the vocabulary problem — there are only ~100 characters to cover — but creates a different one: sequences become extremely long. A 500-word paragraph might be 500 tokens as words, but 2,500+ tokens as characters. Since transformers scale quadratically with sequence length in attention computation, that's a huge cost.

Modern tokenizers split the difference with subword tokenization: common words stay whole, rare words get broken into meaningful chunks, and any string can be represented using a fixed, manageable vocabulary (typically 30,000–100,000+ tokens).

Note

Because rare or unseen words get decomposed into familiar pieces ("unbelievability" → "un" + "believ" + "ability"), the model can still form a reasonable representation of a word it never saw whole during training. This is the same trick that lets it handle misspellings, code identifiers, and words in languages that share subword roots.

Byte Pair Encoding (BPE)

BPE, used by GPT models, starts from individual characters (or bytes) and iteratively merges the most frequent adjacent pair into a new token, repeating until it reaches a target vocabulary size. It's a greedy, frequency-driven compression algorithm originally invented for data compression, repurposed for NLP in 2015.

The result: extremely common sequences ("the", "ing", "tion") become single tokens, while rare words get split into multiple pieces. GPT-family models use a byte-level variant of BPE, which operates on raw UTF-8 bytes rather than characters — this guarantees every possible string, including emoji and unusual Unicode, can always be tokenized without an "unknown token" fallback.

python

WordPiece and SentencePiece

WordPiece, used by BERT, is similar to BPE but chooses merges based on maximizing the likelihood of the training corpus rather than raw frequency — it picks the pair that most increases the probability of the training data when merged, which tends to produce slightly more linguistically coherent subwords.

SentencePiece is a different kind of tool: it's a tokenizer framework (used by T5, LLaMA, and many multilingual models) that treats the input as a raw stream of Unicode characters, including whitespace, rather than pre-splitting on spaces first. This makes it language-agnostic — it works identically well on English, Japanese, or Thai, none of which reliably use whitespace to separate words the way English does. SentencePiece can implement either BPE or a unigram language model as its underlying algorithm.

Tokenizer algorithms compared

AlgorithmUsed byMerge criterionHandles whitespace
BPE (byte-level)GPT-2, GPT-3, GPT-4Most frequent adjacent pairPre-split on whitespace
WordPieceBERT, DistilBERTMaximizes corpus likelihoodPre-split on whitespace
SentencePiece (BPE or unigram)T5, LLaMA, Gemini, mT5Frequency or probability, on raw textTreated as a regular character

Why tokenization determines your API bill

Every major LLM API — OpenAI, Anthropic, Google — prices requests per token, not per word or character. This means tokenization efficiency has direct financial consequences. English text tokenizes at roughly 4 characters per token (about 0.75 tokens per word). But this ratio is not uniform across languages: because tokenizer vocabularies are trained predominantly on English and other high-resource languages, text in languages like Hindi, Thai, or Japanese can require 2–4x more tokens to express the same meaning, making those requests more expensive and eating more of the context window for equivalent content.

Code is another edge case: whitespace-heavy, indentation-sensitive languages like Python tokenize inefficiently under tokenizers built for prose, which is one reason code-specialized tokenizers exist.

Note

A '128K context window' means 128,000 tokens, not words — roughly 90,000–100,000 English words, but potentially far less for dense technical text, code, or non-English languages. This is why the same document can fit comfortably in one language and overflow the context window in another.

Tokenization quirks that explain weird LLM behavior

Several well-known LLM failure modes trace directly back to tokenization. Models are notoriously bad at counting letters in a word (e.g. "how many r's in strawberry") because the model never sees the word as individual characters — it sees one or two subword tokens and has no direct access to the spelling underneath. Similarly, arithmetic on large numbers is unreliable partly because numbers get tokenized inconsistently ("380" might be one token while "381" is split differently), breaking the positional structure a human relies on to do math.

This is also why prompt engineers sometimes insert extra spaces or restructure text — small changes in surface form can shift how a string gets tokenized, subtly changing what the model "sees."

What's next

Tokenization is the entry point; once text becomes tokens, the next step is turning tokens into vectors the model can reason over. See the embeddings lesson for how that happens, or continue to fine-tuning to see how token-level choices affect training cost.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.