← Back to concepts
10 min read

What is a tokenizer?

Large language models cannot read text directly. Every prompt, document, or line of code has to be turned into numbers first, and that conversion is handled by a tokenizer.

What is a tokenizer?

A tokenizer is a program that splits raw text into small chunks called tokens, then looks each token up in a vocabulary — a fixed table that maps every known token to a unique integer ID.

"AI engineering" -> ["AI", " engineer", "ing"] -> [1723, 8842, 291]

The tokenizer passes those integer IDs to the model. The model then looks up an embedding vector for each ID, adds position information, and computes on those vectors. It never sees letters or words directly. Whatever the tokenizer decides to do with the input shapes everything the model computes afterward.

Where tokenization sits in the model input pipeline Raw text enters the tokenizer, which splits it into text fragments. A fixed vocabulary maps each fragment to an integer ID, and the model uses those IDs to select the starting vector representations it processes. Raw text "AI engineering" Tokenizer split text into known pieces Text tokens AI engineer + "ing" Vocabulary AI 1723 engineer 8842 ing 291 Model IDs select starting vectors Human-readable Tokenizer boundary Model computation
The tokenizer is the translation layer between human-readable text and the integer IDs the model is trained to interpret.

Here are a few more examples of how text commonly gets split:

"unbelievable" -> ["un", "believ", "able"]
"ChatGPT" -> ["Chat", "G", "PT"]
"2024" -> ["202", "4"]
"don't" -> ["don", "'t"]
"hello world" -> ["hello", " world"]

Notice that common English words like hello often stay whole, while brand names, numbers, and contractions get split in ways that can look unpredictable at first. That’s a direct result of which fragments were common enough in the training corpus to earn their own token.

Why not just split on words or letters?

A tokenizer needs a strategy for cutting text into pieces, and there are three broad options:

  • Word-level: one token per word. Simple, but the vocabulary explodes with every new name, typo, or made-up word, and anything missing from the vocabulary can’t be represented.
  • Character-level: one token per letter. Nothing is ever missing from the vocabulary, but ordinary sentences turn into very long sequences, which wastes context and compute.
  • Subword: a middle ground. Most modern LLM tokenizers use subword methods such as BPE, often at the byte level.

Subword tokenization builds a vocabulary of frequently occurring fragments from a large body of training text. Common whole words like the or model usually get their own token, while rarer or longer words like tokenization get broken into reusable fragments such as token and ization. This lets the tokenizer represent essentially any input — including misspellings and words it has never seen — while keeping the vocabulary at a manageable, fixed size (often in the tens of thousands of entries).

Word, character, and subword tokenization compared The invented input engineeringhq becomes one unknown token in a word-level vocabulary, thirteen small tokens at character level, or two reusable known fragments with subword tokenization. The comparison shows the tradeoff between vocabulary coverage and sequence length. Same unfamiliar input: "engineeringhq" Word-level [UNKNOWN] Short sequence but the exact word is lost shorter longer Character-level e n g i n e e r i n g h q Every character is covered but the sequence is much longer shorter longer Subword engineering hq Known pieces preserve the text without one token per character shorter longer Vocabulary flexibility increases from left to right; sequence length is minimized by reusable fragments.
Subword tokenization keeps the input recoverable while avoiding both an unlimited word vocabulary and character-by-character sequences.

For example, compare how a subword tokenizer might treat a familiar word versus an invented one:

"engineering" -> ["engineering"]          (1 token - common enough to earn its own entry)
"engineeringhq" -> ["engineering", "hq"]   (2 tokens - a new combination, split into known pieces)
"xqzzytron" -> ["x", "q", "zzy", "tron"]    (4 tokens - nothing like it appeared in training)

The made-up word xqzzytron still gets represented, just less efficiently, falling back toward smaller and smaller fragments the less familiar it looks.

How the vocabulary gets built

Most subword tokenizers are trained before the model itself. The training process scans a huge text corpus and builds larger tokens from smaller pieces.

Byte-Pair Encoding (BPE) starts with small pieces, often bytes or characters, then repeatedly merges the most frequent adjacent pair. If l and o often appear together, the tokenizer may add lo; if lo and w often appear together, it may add low.

WordPiece is similar, but it scores merge candidates by how useful the combined piece is compared with the separate pieces. Unigram starts with many candidate pieces, then removes pieces that are least useful while keeping likely segmentations.

The result is a vocabulary tuned to the statistics of that corpus: common patterns in the training data become single tokens, and less familiar text falls back to smaller pieces.

An illustrative sequence of BPE vocabulary merges A small example corpus containing low, lower, and lowest starts as individual characters. In this simplified BPE example, the frequent pair l plus o is merged into lo, then lo plus w is merged into low. The learned low fragment can then be reused across all three words while suffixes remain separate. Illustrative corpus: "low lower lowest" 1 Start small l o w l o w e r l o w e s t Characters are available as initial pieces 2 Merge l + o lo w lo w e r lo w e s t new: lo 3 Merge lo + w low low e r low e s t new: low Reusable pieces low suffixes Frequent stems become compact Real training uses a huge corpus and many merge steps; the selected pairs depend on the algorithm and data.
In BPE, frequent adjacent pieces earn merged vocabulary entries, so common patterns become shorter token sequences.

This also means tokenizers trained on different data behave differently. A tokenizer trained mostly on English text will often split non-English words, code syntax, or rare technical jargon into more pieces than a tokenizer built with that content in mind. This is also why code often looks “expensive” to tokenize: symbols like {, }, and repeated indentation often consume separate tokens, though some tokenizers merge common whitespace runs. A short snippet of code can use noticeably more tokens than a sentence of plain English with the same character count.

Special tokens and byte fallback

Tokenizers also reserve IDs for special tokens. These tokens are not ordinary words from your prompt. They mark structure that the model was trained to understand:

  • Beginning or end of a sequence.
  • Chat role markers such as system, user, and assistant.
  • Padding tokens used to make a batch the same length.
  • Tool-call, tool-result, image, or audio placeholders in models that support them.

For example, a chat API may turn a simple conversation into a template like this before the model sees it:

Illustrative chat template:
<|system|>
You are a concise assistant.
<|user|>
Summarize this ticket.
<|assistant|>

Those markers are tokens too, and they count toward the context window. They also help the model separate instructions, user text, and its own reply during inference.

Modern byte-level tokenizers rarely need an unknown token. If a word, emoji, filename, or typo is not in the vocabulary as a whole piece, the tokenizer can fall back to smaller byte-level pieces that still reconstruct the original text. The tradeoff is efficiency: unfamiliar text may take more tokens.

Try it yourself

The best way to build intuition for tokenization is to see it happen on real text. OpenAI hosts a free interactive tool for exactly this — paste in a sentence and it will highlight each token in a different color and show you the token count and the underlying IDs.

Try the OpenAI Tokenizer →

A few things worth trying there:

  • Type a common sentence and see how many words map one-to-one to tokens.
  • Try a made-up word, a typo, or a name and watch it get split into smaller pieces.
  • Paste a short block of code and compare its token count to a plain-English sentence of similar length.
  • Try a non-English sentence and notice whether it uses more tokens than the English equivalent for the same idea.

Seeing these splits firsthand makes the rest of this concept path — tokens, context windows, and cost — much more concrete.

For production work, count tokens with the deployed model’s own tokenizer. Token counts differ between model families, and even closely related models may have different chat templates or special-token rules.

Why the tokenizer choice matters

The tokenizer sits at the very front of the pipeline, so its behavior ripples through everything downstream:

  • Vocabulary size is a tradeoff: a larger vocabulary means fewer tokens per input but a bigger, more expensive embedding table for the model to store.
  • Splitting behavior decides how many tokens a given piece of text becomes, which directly affects how much fits in the context window, how long inference takes, and API cost, since usage is billed per token. For example, the sentence “I love AI engineering” might cost 5 tokens, while a dense paragraph of unfamiliar jargon of similar length could easily cost double that.
  • Tokenizer-model pairing is fixed: a model must always run with the exact tokenizer it was trained with. The vocabulary-to-ID mapping is baked in during training, so swapping tokenizers would make the model’s numeric IDs meaningless.
Why a model must use its matching tokenizer In the matching path, tokenizer A maps the text AI to ID 1723 and model A interprets embedding row 1723 as the learned representation for AI. In the mismatched path, tokenizer B maps AI to ID 491, but model A's row 491 was trained for a different token, so the model receives the wrong representation. The integer only has meaning inside one shared vocabulary mapping MATCHED Tokenizer A "AI" -> 1723 Model A row 1723 learned for "AI" Correct starting vector Matches what training taught. SWAPPED Tokenizer B "AI" -> 491 Model A row 491 learned for another token Wrong starting vector Valid ID, wrong learned meaning.
Token IDs are not universal labels: each model learns the meaning of rows defined by its own tokenizer's vocabulary.

Common pitfalls

  • Leading spaces matter: many tokenizers store hello and hello as different pieces.
  • Numbers can split oddly: dates, decimals, and IDs may break into chunks that do not match how people read them.
  • Non-English text can cost more: token efficiency varies by language and tokenizer, especially for languages under-represented in the tokenizer’s training data.
  • Letter counting is hard for LLMs: a model sees token pieces, not a clean character grid, so exact character tasks often need code or tools.

Once text has been tokenized, the model works with numeric tokens and the vectors derived from them — the next concept in this path.