gpt-oss-safeguard-120b
Release in the gpt-oss-safeguard family · version gpt-oss-safeguard-120b
gpt-oss-safeguard-120b is the larger gpt-oss-safeguard model, a text-only safety reasoning model fine-tuned from gpt-oss-120b. The model card gives 117B total parameters with 5.1B active. Its configuration lists 36 layers, 128 experts with 4 used per token, and a maximum of 131,072 position embeddings. The mixture-of-experts weights are quantized to MXFP4; attention, router, embedding, and output layers are not.136
- Model hub: Hugging Face model card (external site: huggingface.co)
- Repository: gpt-oss-safeguard repository (external site: github.com)
- Paper: Technical report (PDF) (external site: cdn.openai.com)
- License: LICENSE (Apache 2.0) (external site: huggingface.co)
Availability and license
Overall availability
Downloadable from Hugging Face without gating.2
Availability is separate from permission: read the license before using or redistributing.
Apache License 2.0 (gpt-oss-safeguard repository documentation and example policy) (external site: github.com)87
The technical report says the models are available under Apache 2.0 and OpenAI's gpt-oss usage policy. The USAGE_POLICY file shipped with the weights asks users to comply with all applicable law and lists no further restrictions. The GitHub repository holds documentation and an example spam policy with a labeled sample set, not model code.657
Component reuse rights
- weights
- Reviewed qualifying license recorded — check scope and conditions
- code
- Unknown — no complete fact-level rights review
- data
- Unknown — no complete fact-level rights review
- documentation
- Reviewed qualifying license recorded — check scope and conditions
No complete system-rights review is recorded for this release.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Safetensors weights in the ungated Hugging Face repository.2 |
| Inference codeIs code for running the model published? | Public | The model card says the model is used like gpt-oss-120b, following the gpt-oss cookbooks. OpenAI's user guide documents running it with vLLM, Hugging Face Transformers, Ollama, and LM Studio.19 |
| Training codeIs the code used to train the model published? | Unknown | No training or fine-tuning code for gpt-oss-safeguard was found in the official repositories reviewed. |
| Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness. | Unknown | The technical report says the models are fine-tunes of gpt-oss trained without any additional biological or cybersecurity data. It does not otherwise describe the post-training data.6 |
| Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable. | Unknown | Not assessed. |
| Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient. | Unknown | Not assessed. |
| Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above. | Unknown | Not assessed. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Unknown | The report says the model was post-trained to reason from a provided policy but gives no further training details.6 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The technical report gives results on internal multi-policy evaluations, OpenAI's 2022 moderation set, ToxicChat, a multilingual benchmark, and chat-safety evaluations. The internal sets and evaluation code were not found to be published.6 |
What it is useful for
Run and use notes
- The model must be used with OpenAI's harmony response format; the model card says it will not work correctly otherwise.1
- OpenAI's user guide places the policy in the system message, where the reasoning effort (low, medium, or high) is also set. It suggests policies of about 400 to 600 tokens and notes small but meaningful accuracy losses as more policies are added.9
- The technical report notes that dedicated classifiers trained on large labeled sets can still outperform the model, and that it can be time- and compute-intensive to run across all platform content.6
Organization context
Provenance and derivatives
Fine-tuned by OpenAI from its gpt-oss-120b open-weight model.16
- Derived from: gpt-oss-120b — Base model named in the Hugging Face card metadata (base_model_relation finetune).
Other releases in the gpt-oss-safeguard family
- gpt-oss-safeguard-20bModel-disclosure tier (USASI rubric v0.2): Open-weight
What this catalog does not know
- Training code: unknown.
- Training-data information: unknown.
- Training-data access: unknown.
- Complete training pipeline: unknown.
- Legacy data assessment (v0.1): unknown.
- Training recipe: unknown.
Have a primary source? How to report a correction.
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.