Transformers are the model architecture behind most modern language AI. They made it practical for models to connect information across long passages and train efficiently on large amounts of data. The high-level flow is covered in LLM architecture; this article focuses on the transformer family and BERT’s encoder branch.
BERT is one of the most important transformer-based models. It helped popularize the pattern of pre-training a general model on large text datasets and then adapting it to focused language tasks.
What is a transformer?
A transformer processes token representations through repeated layers. Each layer usually combines:
- Self-attention, which lets token positions exchange relevant information.
- A feed-forward network, which transforms each token representation.
- Residual connections and normalization, which help information move through the network.
The attention mechanism is the defining part. It lets each token gather information from other tokens instead of relying only on the token immediately before it.
For example, in the sentence:
The trophy did not fit in the suitcase because it was too large.
The word “it” refers to the trophy, not the suitcase. Attention helps the representation for it use information from trophy, even though the words are separated.
Unlike recurrent neural networks, transformers do not need to pass one hidden state through every token in order. During training, they can process many sequence positions in parallel. This made them a strong fit for modern hardware and large datasets.
Encoders and decoders
Transformers can be built with encoders, decoders, or both.
- Encoder models read input and build a representation of its meaning.
- Decoder models generate output one token at a time.
- Encoder-decoder models read an input with an encoder and generate an output with a decoder.
BERT is an encoder-only model. GPT-style models are decoder-only models. Translation models often use encoder-decoder designs.
What is BERT?
BERT stands for Bidirectional Encoder Representations from Transformers. The key word is “bidirectional.” Within an input, BERT’s self-attention can use tokens on both the left and the right when building each contextual representation.
This is different from a decoder-only model that uses a causal mask and generates text from left to right. BERT is not primarily designed to write long answers. It is designed to turn a complete input into representations that another component can classify, compare, or extract information from.
BERT became popular because it performed very well on tasks like:
- Text classification.
- Sentiment analysis.
- Named entity recognition.
- Question answering over a passage.
- Search and ranking.
- Sentence similarity.
BERT’s input format
Original BERT uses WordPiece subword tokenization. A tokenizer can split one word into several tokens, so BERT’s training and task heads work with tokens, not always whole words.
BERT also uses special tokens:
[CLS]appears at the start. Its final output vector is often used for classification.[SEP]separates text segments and marks the end of a segment.[MASK]replaces selected tokens during pre-training.
For example:
[CLS] ticket export fails [SEP] safari only [SEP]
For sentence-pair tasks, the first segment might be a question and the second segment might be a passage. For classification, a task head often reads the final [CLS] representation and turns it into label scores.
How BERT learns
BERT was pre-trained with objectives that force it to use context. Original BERT used masked language modeling plus next sentence prediction. Later variants, such as RoBERTa, dropped next sentence prediction and changed other training details.
In masked language modeling, about 15% of tokens are selected for prediction. Some selected tokens are hidden with [MASK], and the model learns to predict the original token:
The developer fixed the [MASK] before deploying.
To predict a hidden token, the model must use the visible context on both sides. Repeating this across large text datasets helps it learn reusable contextual representations.
Next sentence prediction asked whether two text segments were likely to appear next to each other. It was part of original BERT, but it is not a requirement for every BERT-style encoder.
From pre-training to a task
Pre-training gives BERT reusable contextual representations, but it does not create a finished support-ticket classifier or entity extractor. The model must still be adapted to the target task.
A common approach is fine-tuning:
- Start with a pre-trained BERT encoder.
- Add a small task-specific output layer, often called a task head.
- Train the combined model on labeled examples from the target task.
- Use the adapted model to score new inputs.
For a support-ticket classifier, the training examples might pair each ticket with labels such as billing, technical, or account. BERT produces contextual representations for the ticket. The task head turns those representations into category scores, and fine-tuning adjusts the model so the correct category receives a higher score.
Teams can fine-tune in two common ways. They can update all encoder weights plus the task head, which usually gives better task fit but needs more care. Or they can freeze the encoder and train only the task head, which is cheaper and more stable but less flexible.
Token-level tasks use the representations differently. A named entity recognizer can classify each token as a person, organization, location, or neither. The same pre-trained encoder can therefore support different tasks when it is paired with the right training data and output layer.
BERT vs GPT-style models
BERT and GPT-style models are both transformers, but they are optimized for different jobs.
BERT-style models are strong at focused understanding tasks. They read a complete input and produce representations that can become a label, score, extracted span, or embedding. Because BERT sees both sides of each token, it is good at understanding a complete input, but it is not built to generate text left to right like a causal decoder.
GPT-style models are strong at generation tasks. They produce text step by step and can write explanations, code, summaries, and conversations.
This is a practical distinction, not a strict boundary. A generative model can classify text when prompted, and an encoder can predict masked tokens. The architecture and training objective mainly affect which behavior is natural and efficient.
Simple comparison:
| Model type | Main strength | Common use |
|---|---|---|
| BERT-style encoder | Understanding text | Classification, search, extraction |
| GPT-style decoder | Generating text | Chat, writing, coding, reasoning |
| Encoder-decoder | Transforming input to output | Translation, summarization |
Why BERT still matters
Even though generative models get more attention today, BERT-style models are still useful. They can be smaller, faster, and cheaper for focused understanding tasks.
For example, if you need to classify support tickets into categories, a BERT-style model may be more efficient than calling a large chat model for every ticket.
Search systems also use encoder models to create embeddings or ranking signals. Raw BERT outputs contextual token vectors; a sentence embedding requires pooling, such as using the [CLS] vector or averaging token vectors. Models trained specifically for embeddings usually work much better than raw BERT pooling.
BERT is also a family, not just one model. RoBERTa, DistilBERT, DeBERTa, and domain-specific variants keep the encoder idea but change training data, size, speed, or details of the objective. Encoder models remain common for classification, ranking, extraction, and embeddings because they are often cheaper and easier to serve than large generative models.
When to choose a BERT-style model
A BERT-style encoder can be a strong choice when:
- The output is a label, score, span, ranking, or fixed-size representation.
- You can train or select a model suited to the language and domain.
- Low latency, high throughput, or local deployment matters.
- The task should behave consistently rather than produce open-ended text.
It may be a poor fit when the product needs long free-form responses, multi-step instructions, or flexible conversation. Fine-tuning also requires representative examples. A model trained on general text may perform poorly on specialized legal, medical, or internal language unless it is adapted and evaluated on that domain.
The key idea
Transformers repeatedly combine attention with token-level transformations to build contextual representations. BERT uses an encoder-only design and bidirectional context to understand a complete input. GPT-style models use a causal decoder to generate text one token at a time. Choose a BERT-style encoder when the system needs focused representations, labels, spans, rankings, or embeddings; choose a generative decoder when it needs open-ended text.