gpt-oss-120b
Release in the gpt-oss family · version gpt-oss-120b
gpt-oss-120b is the larger gpt-oss model: a 36-layer mixture-of-experts transformer with 116.8B total and about 5.1B active parameters per token, using 128 experts with the top 4 selected per token. It is text-only, supports context up to 131,072 tokens, and ships with its MoE weights quantized to MXFP4.1
- Model hub: Hugging Face model card (external site: huggingface.co)
- Repository: gpt-oss repository (external site: github.com)
- Paper: Model card (arXiv) (external site: arxiv.org)
- License: LICENSE (Apache 2.0) (external site: huggingface.co)
Availability and license
Overall availability
Downloadable from Hugging Face without gating.2
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | MXFP4-quantized MoE weights plus an original-format checkpoint are in the Hugging Face repository.2 |
| Inference codeIs code for running the model published? | Public | The GitHub repository has reference PyTorch, Triton, and Metal implementations, tools, and a Responses API server.5 |
| Training codeIs the code used to train the model published? | Unknown | The official repository covers inference, tools, and evaluation; no training code was found. |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card describes a text-only dataset of trillions of tokens focused on STEM, coding, and general knowledge, filtered for hazardous biosecurity content, with a June 2024 knowledge cutoff. The data itself is not released.1 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The model card documents architecture, tokenizer, training compute (2.1 million H100-hours), and post-training with chain-of-thought reinforcement learning at a high level, not in reproducible detail.1 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The model card reports benchmark and safety results; the repository includes evaluation code adapted from simple-evals for running GPQA and HealthBench.17 |
What it is useful for
Presented for production, general-purpose, and high-reasoning use cases, including tool use, function calling, and fine-tuning on a single H100 node.2
Run and use notes
- The model card documents running the model with Transformers, vLLM, Ollama, and LM Studio, and OpenAI's repository provides reference PyTorch, Triton, and Metal implementations.25
- OpenAI states that with its MoE weights in MXFP4 precision (4.25 bits per parameter), gpt-oss-120b fits on a single 80GB GPU such as an NVIDIA H100 or AMD MI300X; the model card lists the checkpoint size as 60.8 GiB.12
- The model must be prompted with OpenAI's harmony response format; the model card says it will not work correctly otherwise.2
Organization context
Other releases in the gpt-oss family
- gpt-oss-20bModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.