← Back to concepts
10 min read

Fine-tuning a pre-trained model

A general-purpose model is trained on broad data, then made available for many tasks. Fine-tuning is a second round of training that adapts a pre-trained model using examples from a narrower domain or task.

The goal is not to teach the model one fact for one request. The goal is to change how the model behaves more consistently across many future requests. In a common training path, fine-tuning comes after broad pre-training and post-training, but some teams also fine-tune a base model directly when that fits the use case.

A common path to a fine-tuned model A common training lane shows pre-training on huge mixed data producing a base model, then post-training with supervised fine-tuning and preference tuning or RLHF producing an assistant model, then optional fine-tuning on task or domain examples producing a specialized model. Some projects fine-tune from a base model instead. A lower lane shows prompting and retrieval as runtime inputs that shape one answer but do not update weights. A COMMON PATH - training stages change weights Pre-training huge mixed data, next-token prediction Base model broad patterns, not a product yet Post-training supervised + preference tuning Assistant follows requests Fine-tuning task/domain examples Specialized model changed future behavior CHANGES ONLY THE INPUT - each request Prompt + retrieved context guides one response; weights unchanged runtime input to the deployed model
A common path fine-tunes after post-training, but fine-tuning can also start from a base model; prompts and retrieved context only shape the current request.

What fine-tuning changes

During normal inference, the model’s parameters stay fixed. Prompt engineering can guide the model for one request, but it does not rewrite the model’s weights.

Fine-tuning is different. It runs additional training and updates some model parameters. After fine-tuning, the model itself has changed. It may follow a format more reliably, use domain language more naturally, or handle a repeated task with less instruction in every prompt.

Common fine-tuning goals include:

  • Producing a consistent output format.
  • Adapting tone or style to a specific product.
  • Improving performance on a narrow classification or extraction task.
  • Teaching domain-specific language patterns.
  • Reducing the amount of repeated prompt instruction needed at runtime.

Fine-tuning does not guarantee factual accuracy. It changes behavior, not the need for evaluation.

A small example

Imagine a support team wants every issue summary to follow this shape:

Problem: ...
Likely cause: ...
Next step: ...
Priority: ...

Prompting can ask a general model to use that format. If the task is occasional, that may be enough. But if the product needs thousands of summaries every day and small formatting mistakes break downstream automation, fine-tuning may help.

The training set might include examples like:

Input:
Customer cannot reset password after receiving the reset email.

Output:
Problem: Password reset link does not complete successfully.
Likely cause: Expired link, blocked redirect, or account state issue.
Next step: Ask for timestamp and browser, then generate a fresh reset link.
Priority: Medium

After enough high-quality examples, the fine-tuned model can learn the desired structure and language pattern more reliably than a prompt alone. The important pattern is learned from many examples, not from one perfect demonstration.

How fine-tuning examples teach a repeated behavior Three reviewed issue-summary examples enter a fine-tuning job. The job updates selected model parameters. Later, a new support ticket passes through the tuned model and produces the same four-field structure: Problem, Likely cause, Next step, and Priority. A note says the examples are split into training and held-out evaluation data. REVIEWED EXAMPLES Example 1 Input: password issue Output format: four fields Example 2 Input: billing dispute Output format: four fields Example 3 Input: delivery delay Output format: four fields hold out some examples for evaluation Fine-tuning job continues training on the pattern updates parameters Future production request Input: refund confusion Problem: ... Likely cause: ... Next step: ... Priority: ...
Fine-tuning works best when examples teach a stable pattern that future requests should follow again and again.

Fine-tuning vs prompting vs retrieval

Fine-tuning is often confused with other ways to improve a model. They solve different problems.

Approach What changes Best for
Prompting The temporary input Instructions, examples, formatting, task framing
Retrieval The context supplied at runtime Fresh facts, private documents, source-grounded answers
Fine-tuning Model weights or added adapters Repeated behavior, style, classification, specialized patterns

If the model lacks current or private knowledge, fine-tuning is usually not the first answer. Retrieval and context engineering are often better because the source material can be updated without retraining the model.

If the model misunderstands the task format, start with better prompting and examples. Fine-tuning becomes more attractive when the prompt is large, brittle, expensive, or still unreliable after careful design.

Choosing between prompting, retrieval, and fine-tuning Three problem statements lead to different interventions. If the issue is one request needs clearer instructions, use prompting, which changes only the temporary input. If the issue is missing current or private facts, use retrieval, which supplies context at runtime. If the issue is stable repeated behavior still fails after prompt and retrieval improvements, use fine-tuning, which changes model parameters and requires evaluation. Match the fix to what is actually missing One request needs guidance format, tone, examples, framing Prompting changes input only Missing facts current, private, source-grounded Retrieval changes context only Stable behavior still fails prompt is large, brittle, costly Fine-tuning changes parameters Before fine-tuning define a measured target collect reviewed examples hold out evaluation data compare against baseline do not use it for missing knowledge
Fine-tuning is the right lever only when the desired change is stable model behavior, not a one-off instruction or missing context.

Types of fine-tuning

There are several ways to adapt a model:

  • Full fine-tuning updates many or all model parameters. It can be powerful, but expensive and harder to manage.
  • Parameter-efficient fine-tuning (PEFT) usually trains a small set of added parameters while keeping most of the base model fixed.
  • Instruction fine-tuning trains the model on input-output examples that teach it to follow task instructions.
  • Preference tuning uses comparisons between outputs to encourage preferred responses. It includes RLHF and direct methods such as DPO.

Hosted model providers may hide these details behind a simpler API. Even then, it helps to understand that fine-tuning is training, not configuration.

LoRA and QLoRA adapters

LoRA, short for low-rank adaptation, is a common PEFT method. It freezes the base weights and trains small adapter matrices that are added to some layers. Instead of rewriting a huge weight matrix, LoRA learns a small update that nudges the layer in the right direction.

That adapter file can be much smaller than the base model: often megabytes instead of gigabytes. Teams can swap adapters for different tasks, keep them as separate artifacts, or merge the adapter update into the base weights for deployment. This is especially common with open-weight models where teams control the runtime.

QLoRA uses LoRA while the base model is loaded in a quantized format, often 4-bit, to save memory. The base stays mostly frozen and compressed; the small adapter remains trainable. See parameters and weights for the memory and quantization basics.

LoRA adds a small trainable update to frozen base weights An input vector flows into a large frozen base weight matrix W and also into a small trainable LoRA path. The LoRA path uses a down projection A and an up projection B; their product creates a small update. The base output and LoRA update are added together, so training changes the adapter while the base weights stay frozen. LoRA trains a small add-on path while base weights stay frozen Input x token vector Frozen base weights W large matrix, not trained Base output W x LoRA A down project LoRA B up project Small update B x A x Add W x + update QLoRA keeps the base in a quantized form, such as 4-bit, while the small LoRA adapter is trained.
LoRA changes behavior by training a small adapter update, not by updating every base weight.

Continued pre-training, instruction tuning, and preference tuning

These names are easy to mix up:

  • Continued pre-training uses more raw domain text and the same prediction-style objective. It helps the model absorb domain language patterns, not a specific response format.
  • Instruction fine-tuning uses input-output demonstrations to teach a task, format, or assistant behavior.
  • Preference tuning uses preferred and rejected outputs to move behavior toward what reviewers, rules, or reward signals prefer.

What makes good training data

Fine-tuning is only as good as the examples used to train it. A small, clean dataset often beats a large, noisy one.

Good fine-tuning data should be:

  • Representative: examples match real production inputs.
  • Consistent: similar inputs receive similar outputs.
  • Specific: examples show the behavior you actually want.
  • Reviewed: labels and answers are checked for errors.
  • Deduplicated: repeated examples should not dominate the training signal.
  • Separated from evaluation data: test examples should not be used for training.

Use clear splits: a training set for updates, a validation set for choosing checkpoints and settings, and a final test set for the last comparison. Keep a held-out evaluation set that reflects real inputs, including edge cases. Compare the tuned model against the original model plus a strong prompt baseline, not just against an old weak prompt. Avoid leakage: if test examples or near-duplicates appear in training, the measured gain may be fake.

If the examples contain contradictions, stale policy, private data that should not be learned, or inconsistent formatting, the fine-tuned model can learn those problems too.

Costs and risks

Like other post-training choices, fine-tuning adds operational work. Before choosing it, consider:

  • Data collection: creating examples can take more time than writing prompts.
  • Evaluation: you need tests to prove the tuned model is better.
  • Versioning: each tuned model and adapter is a new artifact to track.
  • Drift: the task, policy, or base model may change later.
  • Overfitting: the model may memorize narrow examples and perform worse on new cases.
  • Catastrophic forgetting: the model may lose general abilities while specializing. Mitigate with mixed general data, lower learning rates, fewer epochs, or adapters.
  • Data privacy: examples can contain sensitive text, so collect and retain them deliberately.
  • Deployment constraints: fine-tuned models may have different cost, latency, or provider support.

Operationally, tune the learning rate, number of epochs, and checkpoint choice on validation results. Keep versioned datasets, prompts, adapters, base model IDs, and evaluation reports so you can roll back when a new tuned model regresses.

Fine-tuning should improve a measured outcome. If you cannot define what better means, it is hard to know whether the tuning helped.

When fine-tuning is worth it

Fine-tuning is most useful when:

  • The task repeats often enough to justify the setup cost.
  • You have high-quality examples of the desired behavior.
  • Prompting and retrieval have been tried and still fall short.
  • The desired behavior is stable, not changing every week.
  • You can evaluate the model before and after tuning.

It is less useful when the problem is missing knowledge, weak product requirements, or poor retrieval. In those cases, changing the model may hide the real issue.

The key idea

Fine-tuning adapts a pre-trained model by continuing training on task-specific examples. It can make repeated behavior more reliable, but it is not a shortcut around good data, prompting, retrieval, or evaluation. Use it when you need a model to consistently behave differently, not just when you need to give it more information.