A machine learning model is a large mathematical function. A large language model is one example. It takes an input, performs many calculations, and produces an output. The values that control those calculations are called parameters. In neural networks, the largest group of parameters is usually weights. A bias is a learned offset added before an activation, which lets a layer shift its output up or down.
You can think of weights as adjustable settings inside the model. They are not rules that a programmer writes one by one. Training changes them so the model becomes better at the task it is learning.
How weights are learned
Training usually starts with weights that contain random or unhelpful values. During pre-training or fine-tuning, the model repeats a learning loop:
- The model receives an example, such as a sequence of tokens.
- It makes a prediction using its current weights.
- A loss function measures how far the prediction is from the expected result.
- Backpropagation computes gradients: signals that show which direction to nudge each weight to reduce the loss.
- An optimizer applies tiny updates using a step size called the learning rate.
This process is repeated across many examples. Backpropagation does not try weights one by one. It computes gradients for all weights efficiently, then the optimizer applies small steps. Over time, those small steps add up to patterns that help the model predict, classify, or transform new inputs.
For a language model, the examples may teach patterns such as which tokens tend to occur together, how code is structured, or how an answer usually follows an instruction. The model does not store these patterns as a simple list of facts. They are distributed across many parameters and interact during each calculation.
A small example
Imagine a simple model that predicts whether a message is about a password problem. It might learn that words such as forgot, reset, and login are useful signals.
At first, the weights connected to those words may be poor, so the model makes many wrong predictions. After seeing labeled examples, training increases or decreases those weights as needed:
Message: "I forgot my password"
Expected label: password problem
The real weights in a modern AI model are far more numerous and the relationships are much more complex. The example is only meant to show that training changes numerical values based on errors; it does not manually write a rule such as “if the message contains forgot, return password problem.”
Parameters, weights, and layers
In everyday discussions, parameters and weights are often used almost interchangeably. More precisely, parameters include every learned value in the model, while weights are the learned values that multiply or connect signals in its layers.
Neural networks are organized into layers. Each layer transforms the values from the previous layer, and its weights determine how strongly different signals influence that transformation. A model with billions of parameters has billions of learned values spread across these layers.
The number of parameters is a rough measure of model capacity, not a complete quality score. In a mixture-of-experts model, there may also be a difference between total parameters and active parameters: the model stores many expert weights, but only routes each token through some of them. Two models with similar parameter counts can behave very differently because of differences in:
- Training data and data quality.
- Model architecture and tokenizer.
- Training objective and optimization process.
- Fine-tuning and alignment.
- The task and domain where the model is used.
Precision, memory, and quantization
A parameter count also helps estimate memory. The rough rule is:
model memory ~= number of parameters x bytes per parameter
For an illustrative 7B-parameter model, 16-bit weights use about 14 GB just for the weights. The same weights in 8-bit form use about 7 GB. In 4-bit form, they use about 3.5 GB. Real serving also needs extra memory for activations, runtime overhead, and the KV cache used during inference, so hardware planning needs a safety margin.
Quantization means storing weights with fewer bits. It can make an open-weight model cheaper or easier to run, and it can improve speed on hardware that supports that format. The tradeoff is that very aggressive quantization can reduce quality, especially on tasks that need careful reasoning or exact formatting.
Parameters are not the same as context
When you call a trained model, its parameters are normally fixed. Your prompt does not rewrite the weights. Instead, the prompt supplies temporary context that changes the output for that request.
For example, a model may know a general pattern for writing a summary. A prompt can ask it to summarize a particular document, but that document is not permanently added to the model’s parameters.
This distinction helps explain several common techniques:
- Prompting supplies instructions and examples for one request.
- Retrieval-augmented generation supplies relevant documents in the context without changing the model.
- Fine-tuning performs additional training that changes all weights or trains small add-on adapters, depending on the method.
What parameter count means in practice
More parameters can give a model room to represent more complex patterns, but bigger is not automatically better. Larger models often require more memory, compute, and time to run. They may also be more expensive to train and serve.
When choosing a model, compare quality on your actual task with:
- Latency: how quickly it responds.
- Cost: how much inference and storage use.
- Memory: whether it fits your hardware or hosting limits.
- Reliability: how consistently it handles real inputs.
- Maintenance: how easy it is to update, evaluate, and operate.
Some techniques reduce the resources needed to store or update parameters. Quantization stores weights with fewer bits, which can reduce memory use. Parameter-efficient fine-tuning updates a smaller set of added parameters instead of changing the whole base model. These methods can improve practical deployment, but they do not remove the need for evaluation.
The key idea
Parameters are the learned numerical values inside a model. Training adjusts them from examples so the model becomes better at its task. A prompt can influence a model’s behavior temporarily, while fine-tuning changes learned weights or adds trained adapters. Understanding that difference helps you choose between prompting, retrieval, fine-tuning, and a different model.