← Back to concepts
10 min read

Transformers and BERT

Transformers are the model architecture behind most modern language AI. They made it practical for models to connect information across long passages and train efficiently on large amounts of data. The high-level flow is covered in LLM architecture; this article focuses on the transformer family and BERT’s encoder branch.

BERT is one of the most important transformer-based models. It helped popularize the pattern of pre-training a general model on large text datasets and then adapting it to focused language tasks.

What is a transformer?

A transformer processes token representations through repeated layers. Each layer usually combines:

  • Self-attention, which lets token positions exchange relevant information.
  • A feed-forward network, which transforms each token representation.
  • Residual connections and normalization, which help information move through the network.

The attention mechanism is the defining part. It lets each token gather information from other tokens instead of relying only on the token immediately before it.

For example, in the sentence:

The trophy did not fit in the suitcase because it was too large.

The word “it” refers to the trophy, not the suitcase. Attention helps the representation for it use information from trophy, even though the words are separated.

Unlike recurrent neural networks, transformers do not need to pass one hidden state through every token in order. During training, they can process many sequence positions in parallel. This made them a strong fit for modern hardware and large datasets.

Encoders and decoders

Transformers can be built with encoders, decoders, or both.

  • Encoder models read input and build a representation of its meaning.
  • Decoder models generate output one token at a time.
  • Encoder-decoder models read an input with an encoder and generate an output with a decoder.

BERT is an encoder-only model. GPT-style models are decoder-only models. Translation models often use encoder-decoder designs.

Where BERT fits in the transformer family Shared transformer building blocks branch into three designs. Encoder-only models such as BERT read a complete input and return contextual representations. Decoder-only GPT-style models use only the available prefix to generate the next token. Encoder-decoder models read a source sequence with an encoder and generate a related sequence with a decoder. Transformer building blocks attention + feed-forward layers Encoder-only BERT complete input visible build contextual representations classify, extract, rank, embed Decoder-only GPT-style models available prefix visible predict the next token chat, write, code, continue Encoder-decoder Read, then generate complete source input visible decoder writes related output translate, summarize, transform
BERT is the encoder-only branch of the transformer family: it turns a complete input into useful representations rather than generating a long sequence token by token.

What is BERT?

BERT stands for Bidirectional Encoder Representations from Transformers. The key word is “bidirectional.” Within an input, BERT’s self-attention can use tokens on both the left and the right when building each contextual representation.

This is different from a decoder-only model that uses a causal mask and generates text from left to right. BERT is not primarily designed to write long answers. It is designed to turn a complete input into representations that another component can classify, compare, or extract information from.

BERT became popular because it performed very well on tasks like:

  • Text classification.
  • Sentiment analysis.
  • Named entity recognition.
  • Question answering over a passage.
  • Search and ranking.
  • Sentence similarity.

BERT’s input format

Original BERT uses WordPiece subword tokenization. A tokenizer can split one word into several tokens, so BERT’s training and task heads work with tokens, not always whole words.

BERT also uses special tokens:

  • [CLS] appears at the start. Its final output vector is often used for classification.
  • [SEP] separates text segments and marks the end of a segment.
  • [MASK] replaces selected tokens during pre-training.

For example:

[CLS] ticket export fails [SEP] safari only [SEP]

For sentence-pair tasks, the first segment might be a question and the second segment might be a passage. For classification, a task head often reads the final [CLS] representation and turns it into label scores.

How BERT learns

BERT was pre-trained with objectives that force it to use context. Original BERT used masked language modeling plus next sentence prediction. Later variants, such as RoBERTa, dropped next sentence prediction and changed other training details.

In masked language modeling, about 15% of tokens are selected for prediction. Some selected tokens are hidden with [MASK], and the model learns to predict the original token:

The developer fixed the [MASK] before deploying.

To predict a hidden token, the model must use the visible context on both sides. Repeating this across large text datasets helps it learn reusable contextual representations.

Next sentence prediction asked whether two text segments were likely to appear next to each other. It was part of original BERT, but it is not a requirement for every BERT-style encoder.

Bidirectional masked-token prediction compared with causal next-token prediction In the BERT panel, the masked position can use both the tokens before it and the tokens after it to predict bug. In the GPT-style panel, the model predicts bug using only the prefix The developer fixed the, while future tokens remain unavailable. The examples simplify the training objectives to show the information each prediction may use. BERT: masked language modeling context is available from both directions The developer fixed the [MASK] before deploying Predict masked token bug GPT-style: causal language modeling only the existing prefix is available The developer fixed the Predict next token bug before deploying future context unavailable causal mask blocks future tokens The token choices are illustrative; production models learn distributions over many possible tokens.
BERT learns to fill a gap using context on both sides, while a causal decoder learns to continue a prefix without seeing future tokens.

From pre-training to a task

Pre-training gives BERT reusable contextual representations, but it does not create a finished support-ticket classifier or entity extractor. The model must still be adapted to the target task.

A common approach is fine-tuning:

  1. Start with a pre-trained BERT encoder.
  2. Add a small task-specific output layer, often called a task head.
  3. Train the combined model on labeled examples from the target task.
  4. Use the adapted model to score new inputs.

For a support-ticket classifier, the training examples might pair each ticket with labels such as billing, technical, or account. BERT produces contextual representations for the ticket. The task head turns those representations into category scores, and fine-tuning adjusts the model so the correct category receives a higher score.

Teams can fine-tune in two common ways. They can update all encoder weights plus the task head, which usually gives better task fit but needs more care. Or they can freeze the encoder and train only the task head, which is cheaper and more stable but less flexible.

Token-level tasks use the representations differently. A named entity recognizer can classify each token as a person, organization, location, or neither. The same pre-trained encoder can therefore support different tasks when it is paired with the right training data and output layer.

From BERT pre-training to a support-ticket classifier The process has three phases. During pre-training, large unlabeled text and masked-token prediction produce a reusable BERT encoder. During fine-tuning, labeled support tickets and a task head adapt that encoder into a classifier. During use, a new ticket enters the fine-tuned classifier and receives a category such as billing. 1. PRE-TRAIN Large unlabeled text books, articles, documents Masked-token objective learn contextual patterns Reusable BERT encoder 2. FINE-TUNE Labeled support tickets billing, technical, account BERT + classification head adapt weights to the task Fine-tuned classifier 3. USE New ticket "I was charged twice" Fine-tuned classifier produce category scores Category: billing
Pre-training creates a reusable encoder; labeled examples and a task head turn it into a model for one specific job.

BERT vs GPT-style models

BERT and GPT-style models are both transformers, but they are optimized for different jobs.

BERT-style models are strong at focused understanding tasks. They read a complete input and produce representations that can become a label, score, extracted span, or embedding. Because BERT sees both sides of each token, it is good at understanding a complete input, but it is not built to generate text left to right like a causal decoder.

GPT-style models are strong at generation tasks. They produce text step by step and can write explanations, code, summaries, and conversations.

This is a practical distinction, not a strict boundary. A generative model can classify text when prompted, and an encoder can predict masked tokens. The architecture and training objective mainly affect which behavior is natural and efficient.

Simple comparison:

Model type Main strength Common use
BERT-style encoder Understanding text Classification, search, extraction
GPT-style decoder Generating text Chat, writing, coding, reasoning
Encoder-decoder Transforming input to output Translation, summarization

Why BERT still matters

Even though generative models get more attention today, BERT-style models are still useful. They can be smaller, faster, and cheaper for focused understanding tasks.

For example, if you need to classify support tickets into categories, a BERT-style model may be more efficient than calling a large chat model for every ticket.

Search systems also use encoder models to create embeddings or ranking signals. Raw BERT outputs contextual token vectors; a sentence embedding requires pooling, such as using the [CLS] vector or averaging token vectors. Models trained specifically for embeddings usually work much better than raw BERT pooling.

BERT is also a family, not just one model. RoBERTa, DistilBERT, DeBERTa, and domain-specific variants keep the encoder idea but change training data, size, speed, or details of the objective. Encoder models remain common for classification, ranking, extraction, and embeddings because they are often cheaper and easier to serve than large generative models.

When to choose a BERT-style model

A BERT-style encoder can be a strong choice when:

  • The output is a label, score, span, ranking, or fixed-size representation.
  • You can train or select a model suited to the language and domain.
  • Low latency, high throughput, or local deployment matters.
  • The task should behave consistently rather than produce open-ended text.

It may be a poor fit when the product needs long free-form responses, multi-step instructions, or flexible conversation. Fine-tuning also requires representative examples. A model trained on general text may perform poorly on specialized legal, medical, or internal language unless it is adapted and evaluated on that domain.

The key idea

Transformers repeatedly combine attention with token-level transformations to build contextual representations. BERT uses an encoder-only design and bidirectional context to understand a complete input. GPT-style models use a causal decoder to generate text one token at a time. Choose a BERT-style encoder when the system needs focused representations, labels, spans, rankings, or embeddings; choose a generative decoder when it needs open-ended text.