Most widely used large language models are built on the transformer architecture. A transformer is a neural network design that processes token representations and repeatedly refines them until the model can make a useful prediction. Some newer models mix in state-space layers or other designs, but transformers remain the main pattern behind today’s chat and coding systems.
The architecture became important because it handles relationships between tokens more directly than older sequence models. It can process many input tokens in parallel during training and can connect tokens that are far apart in a document. A tokenizer prepares the text, embeddings turn token IDs into vectors, and the transformer layers refine those vectors.
The attention mechanism
Transformers use attention to let each token gather information from other tokens in the sequence. For every token, the model calculates mixing weights and combines information from other tokens. Those weights are useful signals, but they are not a reliable explanation of the model’s full reasoning.
Consider this sentence:
The trophy did not fit in the suitcase because it was too large.
To interpret it, the model needs to use the clue too large, which points toward trophy rather than suitcase. Attention gives the token representation a way to use that surrounding evidence, even though the related words are separated by several other tokens. It does not guarantee the model will always choose the right referent.
Attention is often described using three learned projections:
- A query represents what the current token is looking for.
- A key represents what each token can be matched on.
- A value contains the information that can be passed along.
The model compares a query with other keys, turns the scores into weights, and uses those weights to combine the values. The details are mathematical, but the practical idea is simple: tokens exchange information according to learned relevance.
Why position matters
Attention by itself does not tell the model whether a token came first, last, or somewhere in the middle. The sequence needs positional information so that these two inputs are not treated as the same:
The cat chased the dog.
The dog chased the cat.
Transformers add position information to token representations, using a method chosen by the model’s design. This lets the network learn order, distance, and patterns such as which words usually follow others.
Layers stacked together
A transformer is made from many repeated layers. Each layer usually contains:
- An attention block that mixes information between tokens.
- A feed-forward network, a small neural network applied to each token vector separately.
- Residual connections, shortcut paths that carry the earlier signal around a block.
- Normalization, a scaling step that keeps values in a stable range as they move through many layers.
Early layers often focus on local patterns, grammar, or simple relationships. Later layers often combine those signals into more abstract representations of meaning, entities, instructions, and long-range structure. This is a common observed tendency, not a hard rule.
The model does not look up a finished answer in one layer. Each layer makes a small transformation. During generation, the final position’s output representation is converted into scores for possible next tokens.
What sets model size
Model size is mostly set by a few design choices:
- Layer count: more repeated transformer layers mean more transformations and more weights.
- Hidden size: wider token vectors give each token more numeric space to store features.
- Attention heads: more heads let attention run several smaller comparisons in parallel.
- Vocabulary size: a larger token vocabulary needs larger input and output tables.
- Context length: a longer context window does not always add many weights, but it increases memory and compute during use.
Together these choices shape the number of parameters and the cost of serving the model. Some large models use mixture-of-experts feed-forward layers, where only a subset of expert blocks activates for each token. That can increase total parameters without using all of them on every token.
Encoder and decoder designs
Transformer components can be arranged in different ways:
- Encoder-only models read an input and build contextual representations. They are useful for classification, extraction, ranking, and embeddings.
- Decoder-only models generate tokens from left to right. They are common for chat, writing, and code generation.
- Encoder-decoder models use an encoder to read an input and a decoder to generate a related output, such as a translation or summary.
BERT is an encoder-only transformer. GPT-style language models are decoder-only transformers. The underlying building blocks are related, but the attention mask and training objective make them suited to different tasks.
A generation example
Suppose a decoder-only model receives:
The deployment failed because
The input tokens pass through the model’s layers. The final representation at the end of the sequence produces scores for possible next tokens, such as the, a, or the database. The model chooses a token, appends it to the sequence, and runs the process again to generate the next one.
During training, future tokens are present in the training example, but a causal attention mask hides them from each position. During inference, those future tokens do not exist yet. The same left-to-right rule keeps the prediction process aligned with the task: predict the next token using the available prefix.
Engineering tradeoffs
The architecture explains several practical constraints:
- Context length affects memory and compute because attention must process relationships across many tokens.
- Model size depends on factors such as layer count, hidden dimensions, and the number of attention heads. Larger models often cost more to run.
- Latency depends on both the input length and the number of tokens generated.
- Attention variants change serving cost: multi-head attention runs several heads in parallel, grouped-query attention can shrink the KV cache, and sliding-window attention can reduce long-context cost.
- Prompt structure matters because attention does not guarantee that every detail receives equal focus. Clear prompt engineering helps relevant instructions and evidence stand out.
You do not need to implement a transformer to use one well. Understanding its basic flow helps explain why context length, token order, model choice, and generation settings affect an application’s behavior.
The key idea
A transformer turns token representations into richer contextual representations by repeatedly exchanging information through attention and transforming it through stacked layers. Decoder-only transformers use those representations to predict the next token, while encoder-only and encoder-decoder designs support other kinds of understanding and transformation tasks.