Olmo 3.1 32B Think
Release in the OLMo family · version Olmo-3.1-32B-Think
Maintained by Ai2 (Allen Institute for AI)1
A 32B reasoning model from Ai2's December 2025 Olmo 3.1 update. It was produced by continuing the reinforcement learning run behind Olmo 3 32B Think for about three more weeks with additional passes over the Dolci-Think-RL data, and it writes a reasoning trace before its final answer.123
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: open-instruct Olmo 3 post-training scripts (external site: github.com)
- Repository: OLMo-core Olmo 3 training scripts (external site: github.com)
- Paper: Olmo 3 technical report (arXiv 2512.13961) (external site: arxiv.org)
Availability and license
Overall availability
Downloadable from Hugging Face without gating. The card says the model is intended for research and educational use in line with Ai2's Responsible Use Guidelines.1
Availability is separate from permission: read the license before using or redistributing.
Apache License 2.0 (open-instruct) (external site: raw.githubusercontent.com)5
Apache License 2.0 (OLMo-core) (external site: raw.githubusercontent.com)7
The model card states the model is licensed under Apache 2.0 and intended for research and educational use in accordance with Ai2's Responsible Use Guidelines, which list categories of prohibited use. The post-training datasets are ODC-BY but include third-party sources under their own licenses, per their dataset cards.11711
Model-disclosure tier
Every item in the model checklist is public, including the training data itself, and the weights and code are under OSI-approved licenses.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Final weights and RL step checkpoints (step revisions) are published on Hugging Face.1 |
| Inference codeIs code for running the model published? | Public | The model card documents inference with Hugging Face Transformers and vLLM.1 |
| Training codeIs the code used to train the model published? | Public | open-instruct publishes the Olmo 3 32B Think SFT, DPO, and RL scripts, and OLMo-core publishes the 32B base-model pretraining, mid-training, and long-context scripts.46 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Public | The 32B pretraining mix (Dolma 3 Mix 6T), the 32B mid-training and long-context mixes, and the Dolci Think SFT, DPO, and RL datasets used for the 32B Think line are downloadable from Hugging Face without gating under ODC-BY.8910111213 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Public | The technical report describes Olmo 3.1 Think 32B as the Olmo 3 Think 32B RL run continued from 750 to 2,300 steps (an additional 21 days on 224 GPUs). The open-instruct README lists the 32B Think scripts and an experiment report covering 3.1.34 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The model card publishes evaluation results against its SFT, DPO, and Olmo 3 32B Think predecessors; OLMES documents commands for running the Olmo 3 evaluation suites.114 |
What it is useful for
Run and use notes
- The model card documents inference with Hugging Face Transformers 4.57.0 or later and with vLLM, recommending temperature 0.6, top_p 0.95, and up to 32,768 generated tokens.1
Organization context
Provenance and derivatives
Trained by Ai2 from its own Olmo 3 32B base model (Olmo-3-1125-32B) through the Olmo-3-32B-Think-SFT and Olmo-3-32B-Think-DPO checkpoints, then RLVR; Olmo 3.1 extends the Olmo 3 32B Think RL run. The base model was pretrained on the Dolma 3 Mix (6T).13813
- Derived from: Olmo 3 32B (base), Olmo-3-1125-32B (external site: huggingface.co) — Base model; no separate catalog record.
- Derived from: Dolma 3 Mix (6T) — Pretraining data for the 32B base model.
- Derived from: Dolci-Think-RL-32B (external site: huggingface.co) — RL prompts used for Olmo 3 and 3.1 32B Think.
Other releases in the OLMo family
- Olmo 3 7B (base)Model-disclosure tier (USASI rubric v0.1): Open-stack
- Olmo 3 7B InstructModel-disclosure tier (USASI rubric v0.1): Open-stack
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.