Tokenization in NLP
The first step nobody thinks about
Before a language model can do anything with text, it has to convert that text into numbers. Tokenization is the process of breaking raw text into smaller units — tokens — and mapping each one to an integer ID that the model's embedding layer understands.
It sounds like a boring preprocessing detail, but tokenization quietly shapes almost everything about how an LLM behaves: how much you pay per API call, how much text fits in a context window, why the model is bad at counting letters in a word, and why it sometimes stumbles on certain languages or code.
Why not just split on words or characters?
The naive options both fail. Word-level tokenization (splitting on spaces) creates a vocabulary problem: English alone has hundreds of thousands of word forms, and the model would have no way to handle a word it never saw during training — a rare name, a typo, a new slang term all become an unknown token.
Character-level tokenization solves the vocabulary problem — there are only ~100 characters to cover — but creates a different one: sequences become extremely long. A 500-word paragraph might be 500 tokens as words, but 2,500+ tokens as characters. Since transformers scale quadratically with sequence length in attention computation, that's a huge cost.
Modern tokenizers split the difference with subword tokenization: common words stay whole, rare words get broken into meaningful chunks, and any string can be represented using a fixed, manageable vocabulary (typically 30,000–100,000+ tokens).
Note
Byte Pair Encoding (BPE)
BPE, used by GPT models, starts from individual characters (or bytes) and iteratively merges the most frequent adjacent pair into a new token, repeating until it reaches a target vocabulary size. It's a greedy, frequency-driven compression algorithm originally invented for data compression, repurposed for NLP in 2015.
The result: extremely common sequences ("the", "ing", "tion") become single tokens, while rare words get split into multiple pieces. GPT-family models use a byte-level variant of BPE, which operates on raw UTF-8 bytes rather than characters — this guarantees every possible string, including emoji and unusual Unicode, can always be tokenized without an "unknown token" fallback.
WordPiece and SentencePiece
WordPiece, used by BERT, is similar to BPE but chooses merges based on maximizing the likelihood of the training corpus rather than raw frequency — it picks the pair that most increases the probability of the training data when merged, which tends to produce slightly more linguistically coherent subwords.
SentencePiece is a different kind of tool: it's a tokenizer framework (used by T5, LLaMA, and many multilingual models) that treats the input as a raw stream of Unicode characters, including whitespace, rather than pre-splitting on spaces first. This makes it language-agnostic — it works identically well on English, Japanese, or Thai, none of which reliably use whitespace to separate words the way English does. SentencePiece can implement either BPE or a unigram language model as its underlying algorithm.
Tokenizer algorithms compared
| Algorithm | Used by | Merge criterion | Handles whitespace |
|---|---|---|---|
| BPE (byte-level) | GPT-2, GPT-3, GPT-4 | Most frequent adjacent pair | Pre-split on whitespace |
| WordPiece | BERT, DistilBERT | Maximizes corpus likelihood | Pre-split on whitespace |
| SentencePiece (BPE or unigram) | T5, LLaMA, Gemini, mT5 | Frequency or probability, on raw text | Treated as a regular character |
Why tokenization determines your API bill
Every major LLM API — OpenAI, Anthropic, Google — prices requests per token, not per word or character. This means tokenization efficiency has direct financial consequences. English text tokenizes at roughly 4 characters per token (about 0.75 tokens per word). But this ratio is not uniform across languages: because tokenizer vocabularies are trained predominantly on English and other high-resource languages, text in languages like Hindi, Thai, or Japanese can require 2–4x more tokens to express the same meaning, making those requests more expensive and eating more of the context window for equivalent content.
Code is another edge case: whitespace-heavy, indentation-sensitive languages like Python tokenize inefficiently under tokenizers built for prose, which is one reason code-specialized tokenizers exist.
Note
Tokenization quirks that explain weird LLM behavior
Several well-known LLM failure modes trace directly back to tokenization. Models are notoriously bad at counting letters in a word (e.g. "how many r's in strawberry") because the model never sees the word as individual characters — it sees one or two subword tokens and has no direct access to the spelling underneath. Similarly, arithmetic on large numbers is unreliable partly because numbers get tokenized inconsistently ("380" might be one token while "381" is split differently), breaking the positional structure a human relies on to do math.
This is also why prompt engineers sometimes insert extra spaces or restructure text — small changes in surface form can shift how a string gets tokenized, subtly changing what the model "sees."
What's next
Tokenization is the entry point; once text becomes tokens, the next step is turning tokens into vectors the model can reason over. See the embeddings lesson for how that happens, or continue to fine-tuning to see how token-level choices affect training cost.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.