Large language models do not read text as words or sentences. Everything a model processes — the prompt you send, the documents it retrieves, the reply it generates — is first broken down into tokens. The model receives token IDs, turns them into embedding vectors, and computes on those vectors.
What is a token?
A token is a chunk of text produced by a tokenizer: sometimes a whole word, sometimes part of a word, sometimes just punctuation or a single space. There’s no fixed rule like “one token per word” — the split depends on the tokenizer’s vocabulary, which was built from its training data.
"I love AI engineering." -> ["I", " love", " AI", " engineering", "."] (5 tokens)
"unbelievable" -> ["un", "believ", "able"] (3 tokens)
"GPT-4" -> ["G", "PT", "-", "4"] (4 tokens)
Tokenizers may also add special tokens that mark roles, tool calls, images, or the end of a message. Those markers are not visible words in your prompt, but they still count.
Why not just count words or characters?
Words and characters are intuitive units for people, but they don’t match what the model actually processes:
- Counting words undercounts, since many words split into two or more tokens (
tokenization→token+ization). - Counting characters overcounts, since common fragments are compressed into single tokens (
the,ing,tionare often one token each). - Counting tokens is the only number that matches the model’s real internal cost, because tokens are what the model’s context window, compute, and pricing are all measured in.
For example, the sentence “Please summarize this document for me” might look like 6 words, but tokenizes to something closer to 7-8 tokens once punctuation and word-fragments are accounted for.
Why tokens matter in practice
Tokens are the currency of everything an LLM does:
- Context window: every model has a maximum number of tokens it can hold at once, covering the system prompt, conversation history, retrieved documents, and the model’s own output combined. A “128k context window” means 128,000 tokens total, not characters or words.
- Latency: inference generation happens one token at a time, so longer outputs — more tokens — take longer to produce.
- Cost: most API providers price per token, usually per 1,000 or per 1 million tokens, and often charge input and output tokens at different rates.
- Prompt structure: if your system prompt, examples, and retrieved context together use most of the window, there’s less room left for conversation history or a long answer.
Consider a chatbot with a 16k-token context window. If the system prompt and retrieved documents already use 12,000 tokens, only about 4,000 tokens remain for conversation history and the model’s response — which might be just a few paragraphs.
What happens when you exceed the window
A context window is a hard limit on the total tokens the model can process for one request. If the request plus the allowed output is too large, one of two things usually happens:
- The API rejects the request with a context-length error.
- The application silently truncates, summarizes, or drops older history before sending the request.
Many APIs also have a separate max output token setting. That setting caps how long the response may be, but it does not create extra room. A model with a 128k-token context window and a 4k output cap still needs the prompt, retrieved documents, history, hidden template tokens, and response budget to fit inside the total window.
This is why token budgeting is part of product design. You often need to reserve space for the answer, not just squeeze in as many documents as possible.
A worked example
Say you’re building a summarization feature and want to estimate cost before shipping it:
- Input document: ~2,000 words ≈ 2,600 tokens (using the ~4 chars/token rule)
- System prompt and instructions: ~150 tokens
- Expected summary output: ~200 tokens
Total context usage is roughly 2,950 tokens. Billing separates that into 2,750 input tokens and 200 output tokens. If the model charges $3 per million input tokens and $15 per million output tokens, that single summarization costs:
input: 2750 / 1000000 x $3 = $0.00825
output: 200 / 1000000 x $15 = $0.00300
total: $0.01125 (about $0.011)
Multiply by expected daily volume, and token counting turns into a real budgeting exercise, not just a technical curiosity.
Try it yourself
The fastest way to build intuition is to tokenize real examples and see the exact split and count.
Paste in a paragraph you actually plan to send to a model — a support ticket, a code snippet, a prompt template — and compare the reported token count to your rough word-count estimate. Try the same text in a language other than English, too; many tokenizers use noticeably more tokens per word for non-English text since their vocabularies are trained mostly on English data.
Token efficiency varies by language and tokenizer. Languages under-represented in the tokenizer’s training data often need more tokens per word, but the only reliable answer is to measure with the exact tokenizer your model uses.
Measure real usage
Estimates help during design, but production systems should read the usage fields returned by the model API. These fields usually report input tokens, output tokens, and total tokens for the actual request.
Watch the provider’s definitions closely:
- Cached input tokens may be counted or priced separately.
- Some reasoning models count hidden reasoning tokens differently from visible output tokens.
- Multimodal inputs such as images or audio also consume tokens, even when the user did not type text.
- Chat templates and special tokens can add overhead that rough word-count estimates miss.
When designing an LLM feature, estimate token usage for realistic inputs, then measure real API usage and reserve enough context for the model’s response. Token counts are an engineering constraint, not just an implementation detail.