← Back to concepts
6 min read

Post-training AI models

Post-training is the set of learning stages applied after a model has completed broad pre-training. It turns a capable but general base model into something that is more useful, predictable, and aligned with a product’s goals. Alignment means shaping behavior toward intended goals, user needs, and policy limits.

A pre-trained model may be good at continuing text, but it is not reliably optimized to follow a user’s instruction for a concise answer, structured JSON, or a safe refusal. Post-training teaches these interaction patterns, usually before any product-specific fine-tuning is applied. A common first stage is supervised fine-tuning (SFT), where the model learns from demonstrations of good instructions and answers.

Where post-training fits after pre-training A parameter-changing training lane shows pre-training on huge mixed data producing a base model. Post-training then includes supervised fine-tuning with instruction examples and preference tuning or RLHF with comparison data, producing an assistant model. Optional fine-tuning can later specialize the assistant model for a product task. A lower lane shows prompting and retrieved context as runtime input only. CHANGES THE PARAMETERS - training stages Pre-training broad patterns Base model continues text Post-training SFT: instruction data preference / RLHF Assistant useful behavior Fine-tuning optional task CHANGES ONLY THE INPUT - each request Prompt + retrieved context steers inference; weights unchanged solid: training updates weights dashed: runtime input only
Post-training is the bridge between a base model that predicts text and an assistant model that follows instructions and preferences.

Instruction tuning

One common stage is supervised fine-tuning, often called instruction tuning. People or automated systems create examples containing an instruction and a desired response:

Instruction: Summarize the incident in two bullet points.
Response:    - The database timed out during deployment.
             - The team restored service by rolling back.

The model learns to produce responses that resemble the examples. The examples can teach tone, formatting, domain behavior, tool-call syntax, and how to handle unclear requests.

Instruction tuning does not erase pre-training. It continues training the same weights on instruction-response pairs, so it can improve assistant behavior but also cause regressions if the data is narrow or low quality.

How supervised fine-tuning teaches instruction following Reviewed instruction and response examples are split into input tokens and expected response tokens. The instruction goes through the base model to produce a prediction. The model prediction and desired response both feed the supervised loss. The loss sends an update back to the base model weights so future answers resemble the reviewed examples. SUPERVISED FINE-TUNING DATA Instruction Summarize the incident in two bullet points. Desired response - Database timed out - Rollback restored service Base model predicts response Model prediction may miss format or tone Supervised loss compare prediction with the desired response update model weights Instruction tuning continues training the same weights; it can improve behavior while still needing regression checks.
Supervised fine-tuning compares the model's prediction with the desired response, then updates the model weights rather than the prediction.

Preference and safety training

Instruction examples are not enough to describe every response people would prefer. A second stage can compare candidate answers and teach the model which response is more useful, truthful, harmless, or consistent with a policy. This is often built from human feedback, AI feedback, or checkable task outcomes.

Preference training may use:

  • Human rankings of multiple responses.
  • Critiques from trained reviewers.
  • Rules or automated graders.
  • Safety examples and adversarial prompts.

RLHF, or reinforcement learning from human feedback, is one well-known approach: train a reward model from preferences, then use reinforcement learning such as PPO to update the model. DPO, or direct preference optimization, learns directly from preferred and rejected pairs without a separate reward model or PPO loop.

How preference training turns comparisons into behavior A prompt produces three candidate answers. Reviewers compare the answers for helpfulness, truthfulness, safety, and policy fit. The resulting preference data has two clear paths: one arrow feeds a reward model for RLHF, and a separate arrow feeds a direct preference objective. The policy is then updated so future responses are more likely to match the preferred behavior, but evaluation is needed because preference is only a proxy. Prompt same user task Candidate answers A: concise, grounded B: vague, too long C: unsafe detail Reviewer choice A preferred B, C rejected with policy guidance Preference data comparisons, rankings, critiques, safety labels Reward model for RLHF learns to score responses Direct objective uses pairs directly Preference is useful but incomplete: evaluation still checks factuality, safety, and regressions.
Preference data can train a reward model for RLHF or feed a direct preference objective, but both still need independent evaluation.

Common post-training methods

Post-training is not one method. Teams combine several signals depending on the product goal:

  • Supervised fine-tuning (SFT): learns from demonstrations of good instructions and answers; no reward model. Typical use: teach format, tone, tool syntax, and basic assistant behavior.
  • RLHF with PPO: learns from human preference pairs or rankings, which first train a separate reward model. Typical use: optimize helpfulness, safety, and policy fit from comparisons.
  • Direct preference optimization (DPO): learns from preferred and rejected answer pairs; no reward model. Typical use: cheaper preference tuning without a reinforcement-learning loop.
  • RLAIF or constitution feedback: learns from AI judgments, rules, or a written constitution; sometimes uses a reward model. Typical use: scale feedback when human review is limited or policy-heavy.
  • Verifiable-reward reinforcement learning: learns from checkable outcomes such as passing unit tests or correct math answers; no reward model. Typical use: improve code, math, and reasoning tasks with objective rewards.
  • Distillation: learns from outputs or reasoning traces from a stronger model; no reward model. Typical use: train a smaller or cheaper model to imitate useful behavior.

Modern reasoning models often use reinforcement learning with verifiable rewards because the reward can come from a test runner, math checker, or other objective signal. This is different from asking a reviewer which answer sounds better.

What post-training can improve

Post-training commonly improves:

  • Following instructions.
  • Answer structure and tone.
  • Refusing or redirecting unsafe requests.
  • Conversation behavior.
  • Tool and function calling.
  • Performance on a target domain.

Post-training does not add reliable new knowledge by itself. If the model needs fresh or private information, use retrieval and context engineering or update the training data deliberately. Clear prompt engineering can still matter because the post-trained model only responds to what is in the runtime context.

It can also introduce new weaknesses:

  • Capability regressions or forgetting, where a model becomes worse at tasks the base model handled well.
  • Sycophancy, where it agrees with the user too eagerly.
  • Verbosity bias, where longer answers are rewarded even when short answers are better.
  • Over-refusal, where safe requests are blocked because the model learned a broad refusal pattern.
  • Reward hacking, where the model exploits the training signal instead of solving the real task.

Evaluation matters

Teams should evaluate both the base model and the post-trained model during inference. Useful evaluations include:

  • Representative user tasks.
  • Formatting and tool-call checks.
  • Factuality and citation checks.
  • Safety and abuse tests.
  • Regression tests for important capabilities.

It is important to test the model outside the examples used for training. Otherwise, the training set may hide failures. Compare against the base model, a strong prompt-only baseline, and earlier post-trained checkpoints so regressions are visible.

The key idea

Post-training adapts a broadly trained model into a more useful assistant. Supervised fine-tuning teaches desired response patterns, while preference methods, AI feedback, verifiable rewards, and distillation shape quality and behavior. It improves usability, but every optimization should be checked for new failures and tradeoffs.