Inference is what happens when you send input to a trained model and receive output back. Training changes model weights. Inference uses the weights that already exist.
For a language model, inference usually means taking a prompt, processing its tokens, and generating a response one token at a time. The model may be an LLM in a chat product, a code assistant, or a batch job that summarizes documents.
From prompt to prediction
Before the model can answer, the input is converted into tokens. Those tokens move through the model’s layers, where attention and learned weights produce a representation of the current context. That context can include instructions, user messages, retrieved documents, tool results, and earlier output.
At the end of that pass, the model does not directly produce a finished paragraph. It produces logits, raw scores for every token in its vocabulary. A softmax turns those scores into probabilities, which answer a question like:
Given everything so far, what token is most likely to come next?
If the prompt is:
The capital of France is
the model may assign a very high probability to Paris. Other tokens may still have some probability, but they are less likely. The exact distribution depends on the model architecture, weights, prompt, and decoding settings.
Generation is a loop
Text generation repeats the same basic loop:
- Read the current input tokens.
- Compute scores for the next token at the last position.
- Choose one token using a decoding strategy.
- Append that token to the context.
- Repeat until the model stops or reaches a limit.
For example, this illustrative trace shows token-by-token growth, not a full sentence written at once:
Prompt: "Write one sentence about embeddings."
Step 1: "Embeddings"
Step 2: "Embeddings turn"
Step 3: "Embeddings turn data"
Step 4: "Embeddings turn data into"
Step 5: "Embeddings turn data into vectors"
The model is not writing the whole sentence in one operation. It is repeatedly choosing the next token based on the prompt and everything it has already generated. Each decoding step produces scores for the next token at the last position, and one token is chosen from those scores.
Prefill, decode, and the KV cache
Autoregressive inference has two phases.
Prefill processes the prompt tokens. In most transformer models, the prompt tokens can be processed in parallel in one forward pass. This phase drives time to first token, the delay before streaming starts.
Decode generates output tokens. Each new token depends on the tokens before it, so output appears one token at a time. This phase drives tokens per second, the speed of the stream after it starts.
A KV cache makes decode faster. In attention layers, every token creates keys and values. The cache stores those keys and values so the model does not recompute them for earlier tokens at every step. This trades memory for speed. The cache grows with context length and batch size, so very long prompts and many concurrent requests can become memory-bound.
Decoding controls the next token
The model produces probabilities, but the application still has to decide how to pick from them. This choice is called decoding.
Common decoding strategies include:
- Greedy decoding: always pick the highest-probability token. It is simple and repeatable, but it can get stuck in bland or locally good choices.
- Sampling: draw from the probability distribution. Temperature, top-k, and top-p shape which tokens are likely enough to sample.
- Beam search: keep several candidate continuations and score them as they grow. It is common in translation and other encoder-decoder systems, but less common for open-ended chat because it can reduce diversity.
Common decoding settings include:
- Temperature: divides the logits before softmax. Low temperature sharpens the distribution; high temperature flattens it.
- Top-p: limits choices to the smallest group of tokens whose combined probability reaches a threshold.
- Top-k: limits choices to the k most likely tokens.
- Max output tokens: caps how long the response can be.
- Stop sequences: tell the system when to stop generation after a specific pattern appears.
For a customer support classification task, you usually want low randomness. For brainstorming names for a product, you may want more variety.
A worked example
Imagine a model has to continue this prompt:
The server failed because
It might assign top-candidate probabilities like this:
"the" 0.32
"a" 0.18
"of" 0.08
"there" 0.06
"database" 0.04
all other tokens 0.32
Those five named candidates sum to 0.68 because the remaining probability is spread across many other tokens. The temperature diagram below uses the same five named candidates but renormalizes only that small visible set to 1.0, so the numbers are not the full vocabulary distribution.
Temperature divides logits before softmax. With a very low temperature, sampling will likely choose the because the distribution becomes sharp. With greedy decoding, the top token is always chosen regardless of temperature. With a higher temperature, sampling may choose another plausible token. After choosing one token, the model runs again with the longer context.
This is why the same prompt can produce different answers across runs. The model is using probabilities, not a fixed script.
Inference does not change the model
During normal inference, the model’s parameters stay fixed. If a user tells the model a new fact in a prompt, that fact can affect the current response, but inference alone never updates the weights. An application may store conversation history or user memories separately, but that storage is outside the model’s parameters.
This distinction matters:
- Inference uses the trained model to compute an output.
- Prompting supplies temporary instructions and context.
- Retrieval adds external information to the context before inference.
- Fine-tuning changes model parameters through additional training.
If a support bot answers using a policy document placed in the prompt, the model has not learned that policy forever. The application supplied it for that request.
Why inference matters for engineering
Inference is where a model architecture becomes a production system. The same model can feel fast, slow, reliable, or unpredictable depending on how inference is designed.
Important engineering concerns include:
- Latency: track time to first token, total response time, and tokens per second.
- Throughput: batching and continuous batching group active requests so hardware stays busy.
- Cost: most hosted models charge for input and output tokens.
- Streaming: responses can be shown token by token as they are generated.
- Memory: the KV cache, model weights, and batch size all compete for GPU memory.
- Determinism: lower randomness makes outputs more repeatable, but not always perfectly identical across providers or hardware.
- Context limits: the prompt, retrieved documents, conversation history, and output all compete for the same context window.
- Failure handling: applications need timeouts, retries, validation, and fallbacks when output is missing or malformed.
Teams also use inference optimizations. Quantization stores weights with fewer bits to save memory, a tradeoff covered in parameters and weights. Speculative decoding asks a smaller draft model to propose tokens and a larger model to verify them, which can improve speed when the drafts are often right.
For structured tasks, do not rely only on a clever prompt. Use constrained output formats, validators, retries with clear errors, and tests based on real examples.
The key idea
Inference is the runtime process of using a trained model. The model reads tokens, predicts the next token, appends it, and repeats. Prefill, decode, the KV cache, and decoding strategies shape speed and output style, while engineering choices around latency, cost, context, batching, and validation determine whether the model works reliably in a real product.