Pre-training is the first large learning stage for many AI models. The model studies a very large collection of examples and adjusts its parameters so it becomes good at predicting patterns in the data.
For a large language model, those examples may include text, code, and other sequences of tokens. This is usually self-supervised learning: the labels come from the data itself, not from people hand-labeling every example. Pre-training gives the model broad abilities before it is taught how to follow instructions or behave like a product assistant.
What the model learns
During pre-training, a language model learns patterns such as:
- Grammar and sentence structure.
- Which tokens and ideas tend to appear together.
- Common facts and styles of writing.
- Patterns in source code and other structured text.
- Relationships between tokens across different distances.
The model does not receive a neat database of facts. These patterns are distributed across its learned parameters and shaped by the model architecture. The quality of the result depends on the data, objective, and training process.
The prediction loop
GPT-style models usually use causal prediction, also called next-token prediction: predict the next token using only the tokens before it. BERT-style encoder models often use masked-token prediction: hide some tokens inside an input and ask the model to recover them. The Transformers and BERT article explains why those objectives fit different model designs.
A simplified causal language-model training example looks like this:
Input: The server returned an
Expected: error
The model predicts a probability for every possible next token. A loss function compares those probabilities with the expected token. An optimizer then makes a small update to the weights.
This loop runs over many batches:
- Read a batch of token sequences.
- Predict the next token for causal models, or masked tokens for encoder models.
- Calculate the loss.
- Compute gradients showing how changing each parameter would change the loss.
- Update the parameters slightly.
After many updates, the model becomes better at predicting held-out examples and useful patterns in new text.
Data and scale
Pre-training can require enormous datasets and large amounts of compute. Data is usually cleaned, filtered, deduplicated, and mixed from several sources before training. Teams try to remove low-quality, duplicate, unsafe, or private content, but filtering is imperfect.
More data and compute can improve capability, but raw size is not enough. In plain terms, scaling laws say model size, training tokens, and compute should grow in a balanced way. A too-large model trained on too little or low-quality data may waste capacity. A small model trained on carefully chosen data can outperform a larger model on a narrow task.
Scale also creates tradeoffs:
- Low-quality or biased data can teach unwanted patterns.
- Duplicates can make evaluation look better than real generalization.
- Private or copyrighted material may require careful handling.
- Large training runs are expensive and use substantial energy.
The model’s parameter count alone does not tell you how good it will be. Data quality, data mixture, and training design matter just as much.
Evaluation during pre-training
Teams track progress on held-out data: examples that are not used for training. A common metric is validation loss, which measures how surprised the model is by the correct next token. Another common metric is perplexity, which is validation loss expressed as an easier-to-compare score: lower perplexity means the model is less confused by the held-out text.
Evaluation can be misleading if the data is contaminated. Data contamination means benchmark questions, answers, or near-duplicates leaked into the training data. That can inflate scores because the model may have effectively seen the test before. Deduplication and benchmark filtering reduce this risk, but they do not remove it perfectly.
Knowledge cutoff and continued training
A pre-trained model has a knowledge cutoff: it cannot learn events or documents that were not in its training data. Product systems often handle fresh or private information with retrieval and context engineering instead of retraining the whole model.
Some pipelines add later raw-text stages before assistant training. Continued pre-training or mid-training may use higher-quality domain data, code data, math data, or longer-context examples. These stages still teach by prediction; they just focus the data mix before post-training.
Pre-training is not the whole process
Pre-training creates a broadly capable base model, but it may not be a helpful assistant yet. It may continue text well without reliably following a request, returning a chosen format, or refusing unsafe tasks.
Later stages can include:
- Post-training with instruction examples, preference data, and methods such as RLHF.
- Fine-tuning for a specialized task or domain.
- Prompting to guide the fixed model for a particular request.
- Retrieval to supply current or private information in context.
Pre-training changes the model’s general learned parameters. Prompting and retrieval usually change only the input for one request.
The key idea
Pre-training teaches an AI model broad patterns by repeatedly predicting examples and updating its parameters. It provides the foundation for later instruction tuning, fine-tuning, and product behavior, but it does not by itself create a reliable assistant.