Independent project. Not a U.S. government website.

USASI

Explainer · Models, licenses, and openness

Pretraining, fine-tuning, and post-training

How does a raw model become an assistant, and what does it mean when one model is built on another?

Intermediate6 min readReviewed Oct 8, 2026

General information, not legal or professional advice. All explainers

How this page was made
  • Researched and written with AI assistance from primary sources, which are listed at the end with the date they were read.
  • Source-checked: a separate AI fact-check pass compared each sentence with its source and corrected what did not match (fact-check report (external site: github.com)).
  • Automated checks passed: links, structure, and formatting are validated before every publish.
  • Not individually reviewed by a person before publication. What these review levels mean

Key takeaways

  • A base model is pretrained to predict the next token; post-training stages such as SFT, preference tuning, and RLVR then shape how it responds.
  • The Llama 3.1 license sets conditions on distributed derivatives, such as a name starting with "Llama"; Apache 2.0 lets you license your changes differently.
  • A card's base_model field may name only the previous stage, so follow the chain and confirm the base in the technical report.
On this page

A language model starts as a base model, trained on a very large amount of text to predict what comes next. That makes it good at continuing text but not tuned to follow instructions. Developers then post-train it in stages, such as supervised fine-tuning on example answers, preference tuning on pairs of better and worse answers, and reinforcement learning on tasks whose answers a program can check, and release the results under names such as "Instruct," "Chat," or "Think." Anyone with a base model's weights can post-train it further; whether the base model's license follows the result depends on that license. A model card's base_model field and the developer's technical report tell you what a model was built on. This page is general information, not legal advice.

Pretraining and the base model

A model's weights, also called parameters, are the numbers it learns in training; "8B" means about 8 billion of them. In pretraining, the model reads text split into tokens (words or pieces of words) and adjusts its weights to predict each next token better. OpenAI's InstructGPT paper (external site: arxiv.org) describes this objective as predicting the next token on a webpage, and notes that it differs from following a user's instructions helpfully and safely.

Meta's Llama 3.1 model card (external site: github.com) says those models were pretrained on about 15 trillion tokens from publicly available sources. Ai2's Olmo 3 report (external site: arxiv.org) splits base-model training into pretraining on up to 5.9 trillion tokens, "midtraining" on 100 billion tokens of specialized data, and a long-context extension. Meta's card says pretrained models can be adapted for a variety of text-generation tasks, while its instruction-tuned models are intended for assistant-like chat.

Post-training, stage by stage

Each post-training stage is a form of fine-tuning: continued training on new data, starting from the previous stage's weights. The InstructGPT paper set out three steps: fine-tune on demonstrations written by human labelers; train a separate reward model to predict which of two outputs labelers prefer; then use a reinforcement-learning algorithm, PPO, to push the model toward outputs the reward model scores highly. The last two steps are reinforcement learning from human feedback (RLHF).

Supervised fine-tuning

Supervised fine-tuning (SFT), or instruction tuning, trains on examples of the desired response. Hugging Face's TRL documentation (external site: huggingface.co) shows a conversation as a list of messages: the user asks "What color is the sky?" and the assistant answers "It is blue."

Preference tuning: RLHF and DPO

Preference tuning trains on comparisons. The DPO paper (external site: arxiv.org) describes RLHF as complex and often unstable, and introduces Direct Preference Optimization, which reaches the same goal with a simple classification loss and no sampling from the model during training. TRL's DPO trainer (external site: huggingface.co) expects records with a prompt, a chosen answer, and a rejected one:

{"prompt": "The sky is", "chosen": " blue.", "rejected": " green."}

Reinforcement learning with verifiable rewards

Ai2's Tülu 3 report (external site: arxiv.org) introduced reinforcement learning with verifiable rewards (RLVR). Instead of a reward model, it uses tasks with checkable outcomes, such as math problems and precise instruction following, and gives a reward only when the model's answer is verified as correct. Tülu 3's recipe runs SFT, then DPO, then RLVR.

Instruct, chat, and think variants

Names describe what a model was tuned for. They are conventions, not standards, so read the card.

  • Chat. Meta's Llama 2 model card (external site: github.com) calls its fine-tuned models Llama-2-Chat, optimized for dialogue.
  • Instruct. Llama 3.1 came as pretrained and instruction-tuned models, and its card says the tuned versions use SFT and RLHF.
  • Think. Ai2 post-trained the Olmo 3 base into Olmo 3 Think, trained with SFT, DPO, and RLVR to write a structured thinking trace before answering, and Olmo 3 Instruct, which answers without one and is tuned for function calling.
  • Hybrid. Nous Research's Hermes 4 70B card (external site: huggingface.co) says the model emits a <think>…</think> segment when it decides to deliberate, and a chat-template flag or system prompt turns reasoning mode on.

When one model is built on another

Some developers post-train another developer's base model; others run every stage themselves.

What carries over: licenses

The Llama 3.1 Community License (external site: github.com) lets licensees create derivative works, with conditions. Anyone who distributes or makes available the Llama materials, a derivative of them, or a product or service that contains them (including another AI model) must provide a copy of the agreement and prominently display "Built with Llama." Copies of the Llama materials must also carry Meta's attribution notice in a "Notice" text file. Anyone who uses the materials or their outputs to create, train, fine-tune, or otherwise improve an AI model that is distributed or made available must "include “Llama” at the beginning of any such AI model name." Use must also follow Meta's acceptable use policy, which the license incorporates.

Ai2's Tülu 3.1 card says all Llama 3.1 Tülu 3 models are released under that agreement, and that the Gemma Terms of Use and the Qwen License Agreement also apply because the training mix included outputs from third-party models. Post-training data can bring terms of its own. The Hermes 4 70B card lists llama3 in its license field, although its base_model is Llama 3.1 70B.

The Apache License 2.0 (external site: apache.org) works differently. Section 4 allows distributing derivative works if you include the license, mark files you changed, and keep the original notices, and it lets you put your own modifications, or the derivative as a whole, under additional or different terms, provided you still comply with it. The Olmo 3 cards use Apache 2.0 and say the models are intended for research and educational use under Ai2's Responsible Use Guidelines.

How to find out what a model was built on

Start with the base_model field in the card's metadata. Hugging Face's model card documentation (external site: huggingface.co) says you can set it for a fine-tune, an adapter, a quantized version, or a merge, and the Hub lets you filter models by it. Hermes 4 70B lists meta-llama/Meta-Llama-3.1-70B.

The field may name only the previous step: Olmo 3 7B Instruct lists allenai/Olmo-3-7B-Instruct-DPO, its own DPO checkpoint, so follow the chain. The field is optional and set by the uploader, so confirm it in the technical report. If neither says, treat the base as unknown.

Worked example: tracing a model's lineage

Riley is a fictional reader invented for this page. Riley is a graduate student who wants to fine-tune a small open model for a class project and publish the weights.

  1. Open the cards. Riley compares Llama 3.1 Tülu 3.1 8B and Olmo 3 7B Instruct. Each card's base_model points to a DPO checkpoint.
  2. Follow the chain. The stage tables lead back to Meta's Llama 3.1 8B for Tülu and to Ai2's own Olmo 3 7B base for Olmo.
  3. Read the base licenses. A Tülu-based release would follow the Llama 3.1 agreement (a copy of it, the attribution notice, "Built with Llama," a name starting with "Llama," and the acceptable use policy) plus the Gemma and Qwen terms Ai2 lists. An Olmo-based release would need the Apache 2.0 notices, and Riley could license the changes differently.
  4. Pick a starting point. Because Ai2 publishes each stage, Riley could start from the SFT checkpoint and add a preference-tuning step.

Riley's outcome: start from the Olmo 3 7B Instruct SFT checkpoint and state its lineage on the new card. A student who needs something only Tülu offers could choose Tülu and follow its terms.

What you can do next

Sources

All read on October 8, 2026.

Check your understanding

Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.

  1. 1.What is a base model good at after pretraining, according to the page?
  2. 2.Which post-training method skips a reward model and gives a reward only when the model's answer is verified as correct?
  3. 3.Does a base model's license follow a model that someone post-trains from it?

0 of 3 answered.

Keep learning

Part of Building with AI.

Support Us

Help keep USASI useful.

Find the catalog useful? Leave an optional tip to support its upkeep. Tips never affect listings, coverage, or openness assessments.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project