Llama 3.1 Tülu 3.1 8B
Release in the Tülu family · version Llama-3.1-Tulu-3.1-8B
Maintained by Ai2 (Allen Institute for AI)1
An 8B instruction-following model from Ai2's Tülu 3 line, post-trained from Meta's Llama 3.1 8B base model through supervised fine-tuning and DPO, then reinforcement learning with verifiable rewards. Version 3.1 differs from the original Tülu 3 8B only in the final RL stage, which switched from PPO to GRPO (without a reward model) with further hyperparameter tuning.124
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: open-instruct (training code) (external site: github.com)
- Repository: OLMES (evaluation code) (external site: github.com)
- Paper: Tülu 3 technical report (arXiv 2411.15124) (external site: arxiv.org)
Availability and license
Overall availability
Downloadable from Hugging Face without gating, under the Llama 3.1 Community License Agreement.1
Availability is separate from permission: read the license before using or redistributing.
Llama 3.1 Community License Agreement (external site: raw.githubusercontent.com)16
Apache License 2.0 (open-instruct) (external site: raw.githubusercontent.com)5
The weights inherit Meta's Llama 3.1 Community License, a custom license that requires displaying "Built with Llama", starting derived model names with "Llama", following Meta's Acceptable Use Policy, and a separate license from Meta for licensees whose products had more than 700 million monthly active users on the Llama 3.1 release date. The Tülu 3 SFT mixture is ODC-BY but includes subsets under other licenses, some non-commercial, and outputs of third-party models.623
Model-disclosure tier
Open-weight, plus published inference code, training code, and training recipe, and at least documented training-data composition.
The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Final weights and the preceding SFT and DPO checkpoints are on Hugging Face.12 |
| Inference codeIs code for running the model published? | Public | The card documents loading with Hugging Face Transformers and serving with vLLM.1 |
| Training codeIs the code used to train the model published? | Public | Training code is in allenai/open-instruct; the card gives the exact commit and command used for the 3.1 GRPO run.15 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The post-training data (SFT mixture, 8B preference mixture, and RLVR prompts) is public on Hugging Face, but the card states that the size and composition of the Llama 3.1 pretraining corpus is unknown.1234 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Public | The technical report documents the SFT, DPO, and RLVR recipe; the card lists the GRPO hyperparameters and the final checkpoint step for version 3.1.41 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The card links Ai2's OLMES evaluation repository and shows evaluation scores over training steps; the report describes the Tülu 3 evaluation suite and decontamination.14 |
What it is useful for
The card describes the Tülu 3 models as aimed at chat as well as tasks such as math (MATH, GSM8K) and precise instruction following (IFEval). It notes the models have limited safety training and no built-in response filtering.1
Run and use notes
- The card documents serving with vllm serve allenai/Llama-3.1-Tulu-3.1-8B and suggests --max_model_len=8192 because of the long Llama chat template.1
Organization context
Provenance and derivatives
Fine-tuned by Ai2 from Meta's Llama 3.1 8B base model: first to Llama-3.1-Tulu-3-8B-SFT on the Tülu 3 SFT mixture, then to Llama-3.1-Tulu-3-8B-DPO on the 8B preference mixture, then GRPO on the RLVR-GSM-MATH-IF-Mixed-Constraints prompts. Llama 3.1 was developed by Meta, not by Ai2.124
- Derived from: Llama 3.1 8B (meta-llama/Llama-3.1-8B) — Base model from Meta; catalog link is to the Llama family record.
- Derived from: Llama-3.1-Tulu-3-8B-DPO (external site: huggingface.co) — Immediate parent checkpoint.
- Derived from: Tülu 3 SFT mixture (external site: huggingface.co) — Supervised fine-tuning data (ODC-BY).
Other releases in the Tülu family
No other releases in this family have been assessed.
U.S. eligibility
Eligible · basis: U.S. nonprofit or lab
Post-trained and released by Ai2 (Allen Institute for AI), which describes itself as a Seattle-based non-profit AI research institute. The base model, Llama 3.1 8B, was developed by Meta and is recorded as provenance; Ai2's post-training does not make the base model an Ai2 model.127
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.