USASI
Model release

Llama 3.1 Tülu 3.1 8B

Release in the Tülu family · version Llama-3.1-Tulu-3.1-8B

Maintained by Ai2 (Allen Institute for AI)1

An 8B instruction-following model from Ai2's Tülu 3 line, post-trained from Meta's Llama 3.1 8B base model through supervised fine-tuning and DPO, then reinforcement learning with verifiable rewards. Version 3.1 differs from the original Tülu 3 8B only in the final RL stage, which switched from PPO to GRPO (without a reward model) with further hyperparameter tuning.124

Last reviewedEntry updated Documented release Unknown

Availability and license

Overall availability

Public

Downloadable from Hugging Face without gating, under the Llama 3.1 Community License Agreement.1

Availability is separate from permission: read the license before using or redistributing.

The weights inherit Meta's Llama 3.1 Community License, a custom license that requires displaying "Built with Llama", starting derived model names with "Llama", following Meta's Acceptable Use Policy, and a separate license from Meta for licensees whose products had more than 700 million monthly active users on the Llama 3.1 release date. The Tülu 3 SFT mixture is ODC-BY but includes subsets under other licenses, some non-commercial, and outputs of third-party models.623

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.1. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.1): Open-stack

Open-weight, plus published inference code, training code, and training recipe, and at least documented training-data composition.

The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for Llama 3.1 Tülu 3.1 8B
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicFinal weights and the preceding SFT and DPO checkpoints are on Hugging Face.12
Inference codeIs code for running the model published?PublicThe card documents loading with Hugging Face Transformers and serving with vLLM.1
Training codeIs the code used to train the model published?PublicTraining code is in allenai/open-instruct; the card gives the exact commit and command used for the 3.1 GRPO run.15
Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access.PartialThe post-training data (SFT mixture, 8B preference mixture, and RLVR prompts) is public on Hugging Face, but the card states that the size and composition of the Llama 3.1 pretraining corpus is unknown.1234
Training recipeAre the training configuration and procedure documented in enough detail to follow?PublicThe technical report documents the SFT, DPO, and RLVR recipe; the card lists the GRPO hyperparameters and the final checkpoint step for version 3.1.41
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PublicThe card links Ai2's OLMES evaluation repository and shows evaluation scores over training steps; the report describes the Tülu 3 evaluation suite and decontamination.14

What it is useful for

The card describes the Tülu 3 models as aimed at chat as well as tasks such as math (MATH, GSM8K) and precise instruction following (IFEval). It notes the models have limited safety training and no built-in response filtering.1

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The card documents serving with vllm serve allenai/Llama-3.1-Tulu-3.1-8B and suggests --max_model_len=8192 because of the long Llama chat template.1

Organization context

Provenance and derivatives

Fine-tuned by Ai2 from Meta's Llama 3.1 8B base model: first to Llama-3.1-Tulu-3-8B-SFT on the Tülu 3 SFT mixture, then to Llama-3.1-Tulu-3-8B-DPO on the 8B preference mixture, then GRPO on the RLVR-GSM-MATH-IF-Mixed-Constraints prompts. Llama 3.1 was developed by Meta, not by Ai2.124

Other releases in the Tülu family

No other releases in this family have been assessed.

Tülu family overview

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. nonprofit or lab

Post-trained and released by Ai2 (Allen Institute for AI), which describes itself as a Seattle-based non-profit AI research institute. The base model, Llama 3.1 8B, was developed by Meta and is recorded as provenance; Ai2's post-training does not make the base model an Ai2 model.127

Assessed Sep 29, 2026

Sources

  1. 1.
  2. 2.
  3. 3.
  4. 4.
  5. 5.
  6. 6.
    Llama 3.1 Community License Agreement (external site: raw.githubusercontent.com)

    Meta · License · published Jul 23, 2024 · accessed Sep 29, 2026

  7. 7.
    About us | Ai2 (external site: allenai.org)

    Ai2 · Official page · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project