← Back to concepts
7 min read

Pre-training AI models

Pre-training is the first large learning stage for many AI models. The model studies a very large collection of examples and adjusts its parameters so it becomes good at predicting patterns in the data.

For a large language model, those examples may include text, code, and other sequences of tokens. This is usually self-supervised learning: the labels come from the data itself, not from people hand-labeling every example. Pre-training gives the model broad abilities before it is taught how to follow instructions or behave like a product assistant.

What the model learns

During pre-training, a language model learns patterns such as:

  • Grammar and sentence structure.
  • Which tokens and ideas tend to appear together.
  • Common facts and styles of writing.
  • Patterns in source code and other structured text.
  • Relationships between tokens across different distances.

The model does not receive a neat database of facts. These patterns are distributed across its learned parameters and shaped by the model architecture. The quality of the result depends on the data, objective, and training process.

The prediction loop

GPT-style models usually use causal prediction, also called next-token prediction: predict the next token using only the tokens before it. BERT-style encoder models often use masked-token prediction: hide some tokens inside an input and ask the model to recover them. The Transformers and BERT article explains why those objectives fit different model designs.

A simplified causal language-model training example looks like this:

Input:    The server returned an
Expected: error

The model predicts a probability for every possible next token. A loss function compares those probabilities with the expected token. An optimizer then makes a small update to the weights.

How one sentence becomes next-token training examples The sentence The server returned an error yields four training examples by shifting through it: The predicts server, The server predicts returned, The server returned predicts an, and The server returned an predicts error. For the last example, the model assigns illustrative probabilities of 0.41 to error, 0.18 to empty, 0.12 to invalid, 0.09 to unexpected, and 0.20 to all other tokens. The loss is minus log of 0.41, about 0.89. If the probability of error rose to 0.90, the loss would fall to about 0.11. LABELS COME FROM THE TEXT ITSELF context (input) expected next The server The server returned The server returned an The server returned an error No human labels needed: every position in the text is a prediction exercise. PREDICTION FOR THE LAST EXAMPLE Predicted next-token probabilities error 0.41 expected empty 0.18 invalid 0.12 unexpected 0.09 all others 0.20 loss = -log(0.41) = 0.89 If p(error) rose to 0.90, loss would drop to 0.11. Probabilities are illustrative and sum to 1.
Pre-training turns raw text into its own answer key, then rewards the model for putting more probability on the token that actually came next.

This loop runs over many batches:

  1. Read a batch of token sequences.
  2. Predict the next token for causal models, or masked tokens for encoder models.
  3. Calculate the loss.
  4. Compute gradients showing how changing each parameter would change the loss.
  5. Update the parameters slightly.

After many updates, the model becomes better at predicting held-out examples and useful patterns in new text.

Data and scale

Pre-training can require enormous datasets and large amounts of compute. Data is usually cleaned, filtered, deduplicated, and mixed from several sources before training. Teams try to remove low-quality, duplicate, unsafe, or private content, but filtering is imperfect.

Preparing pre-training data Web text, code, and books and papers flow into a cleaning step that fixes encoding and strips markup. A filtering step tries to drop low-quality, unsafe, or private content. A deduplication step removes repeated documents. A mixing step sets the share of each source. The result becomes tokenized training batches for the pre-training loop. SOURCES web text code books + papers Clean fix encoding, strip markup Filter try to drop low quality, unsafe Deduplicate remove repeated documents Mix set the share of each source Tokenized training batches -> pre-training loop
What survives cleaning, filtering, deduplication, and mixing is what the model learns from; exact steps and order vary between projects.

More data and compute can improve capability, but raw size is not enough. In plain terms, scaling laws say model size, training tokens, and compute should grow in a balanced way. A too-large model trained on too little or low-quality data may waste capacity. A small model trained on carefully chosen data can outperform a larger model on a narrow task.

Scale also creates tradeoffs:

  • Low-quality or biased data can teach unwanted patterns.
  • Duplicates can make evaluation look better than real generalization.
  • Private or copyrighted material may require careful handling.
  • Large training runs are expensive and use substantial energy.

The model’s parameter count alone does not tell you how good it will be. Data quality, data mixture, and training design matter just as much.

Evaluation during pre-training

Teams track progress on held-out data: examples that are not used for training. A common metric is validation loss, which measures how surprised the model is by the correct next token. Another common metric is perplexity, which is validation loss expressed as an easier-to-compare score: lower perplexity means the model is less confused by the held-out text.

Evaluation can be misleading if the data is contaminated. Data contamination means benchmark questions, answers, or near-duplicates leaked into the training data. That can inflate scores because the model may have effectively seen the test before. Deduplication and benchmark filtering reduce this risk, but they do not remove it perfectly.

Knowledge cutoff and continued training

A pre-trained model has a knowledge cutoff: it cannot learn events or documents that were not in its training data. Product systems often handle fresh or private information with retrieval and context engineering instead of retraining the whole model.

Some pipelines add later raw-text stages before assistant training. Continued pre-training or mid-training may use higher-quality domain data, code data, math data, or longer-context examples. These stages still teach by prediction; they just focus the data mix before post-training.

Pre-training is not the whole process

Pre-training creates a broadly capable base model, but it may not be a helpful assistant yet. It may continue text well without reliably following a request, returning a chosen format, or refusing unsafe tasks.

Later stages can include:

  • Post-training with instruction examples, preference data, and methods such as RLHF.
  • Fine-tuning for a specialized task or domain.
  • Prompting to guide the fixed model for a particular request.
  • Retrieval to supply current or private information in context.

Pre-training changes the model’s general learned parameters. Prompting and retrieval usually change only the input for one request.

A common path from pre-training to later stages A common path is shown. In the top lane, training stages change the parameters: pre-training on a huge mixed dataset produces a base model that continues text well. Post-training with instruction examples and preference data produces an assistant model that follows requests. Optional fine-tuning specializes it for a task. In the bottom lane, each request supplies a prompt and retrieved documents to the deployed model. This shapes one answer but leaves the weights unchanged. A COMMON PATH - training stages change weights Pre-training huge mixed data, next-token prediction Base model continues text well Post-training instruction examples, preference data Assistant model follows requests Fine-tuning optional: adapt to a task CHANGES ONLY THE INPUT - each request Prompt + retrieved documents shape one answer; weights unchanged sent to the deployed model solid: training that updates weights dashed: temporary input only
A common path is pre-training, post-training, and optional fine-tuning; prompts and retrieval steer one request without changing weights.

The key idea

Pre-training teaches an AI model broad patterns by repeatedly predicting examples and updating its parameters. It provides the foundation for later instruction tuning, fine-tuning, and product behavior, but it does not by itself create a reliable assistant.