Independent project. Not a U.S. government website.

USASI
Model release

gpt-oss-safeguard-120b

Release in the gpt-oss-safeguard family · version gpt-oss-safeguard-120b

Maintained by OpenAI61

gpt-oss-safeguard-120b is the larger gpt-oss-safeguard model, a text-only safety reasoning model fine-tuned from gpt-oss-120b. The model card gives 117B total parameters with 5.1B active. Its configuration lists 36 layers, 128 experts with 4 used per token, and a maximum of 131,072 position embeddings. The mixture-of-experts weights are quantized to MXFP4; attention, router, embedding, and output layers are not.136Fact reviewed Oct 1, 2026

Last reviewedEntry updated Documented release Oct 2025

Availability and license

Overall availability

Public

Downloadable from Hugging Face without gating.2

Availability fact review: Oct 1, 2026

Availability is separate from permission: read the license before using or redistributing.

The technical report says the models are available under Apache 2.0 and OpenAI's gpt-oss usage policy. The USAGE_POLICY file shipped with the weights asks users to comply with all applicable law and lists no further restrictions. The GitHub repository holds documentation and an example spam policy with a labeled sample set, not model code.657Fact reviewed Oct 1, 2026

Component reuse rights

A readable or downloadable component is not automatically reusable. These indicators concern recorded license evidence, not system certification.
weights
Reviewed qualifying license recorded — check scope and conditions
code
Unknown — no complete fact-level rights review
data
Unknown — no complete fact-level rights review
documentation
Reviewed qualifying license recorded — check scope and conditions

No complete system-rights review is recorded for this release.

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.2. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.2): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.2. Unknown means unassessed or insufficient evidence.
Public materials checklist for gpt-oss-safeguard-120b
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicSafetensors weights in the ungated Hugging Face repository.2
Inference codeIs code for running the model published?PublicThe model card says the model is used like gpt-oss-120b, following the gpt-oss cookbooks. OpenAI's user guide documents running it with vLLM, Hugging Face Transformers, Ollama, and LM Studio.19
Training codeIs the code used to train the model published?UnknownNo training or fine-tuning code for gpt-oss-safeguard was found in the official repositories reviewed.
Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness.UnknownThe technical report says the models are fine-tunes of gpt-oss trained without any additional biological or cybersecurity data. It does not otherwise describe the post-training data.6
Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable.UnknownNot assessed.
Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient.UnknownNot assessed.
Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above.UnknownNot assessed.
Training recipeAre the training configuration and procedure documented in enough detail to follow?UnknownThe report says the model was post-trained to reason from a provided policy but gives no further training details.6
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PartialThe technical report gives results on internal multi-policy evaluations, OpenAI's 2022 moderation set, ToxicChat, a multilingual benchmark, and chat-safety evaluations. The internal sets and evaluation code were not found to be published.6

What it is useful for

Classifying text against a safety policy written by the user, for example for filtering LLM inputs and outputs or labeling content, with the model's reasoning visible to developers. OpenAI's README presents the 120b size for production, general-purpose, high-reasoning use.17Fact reviewed Oct 1, 2026

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The model must be used with OpenAI's harmony response format; the model card says it will not work correctly otherwise.1
  • OpenAI's user guide places the policy in the system message, where the reasoning effort (low, medium, or high) is also set. It suggests policies of about 400 to 600 tokens and notes small but meaningful accuracy losses as more policies are added.9
  • The technical report notes that dedicated classifiers trained on large labeled sets can still outperform the model, and that it can be time- and compute-intensive to run across all platform content.6

Organization context

Provenance and derivatives

Fine-tuned by OpenAI from its gpt-oss-120b open-weight model.16

  • Derived from: gpt-oss-120b — Base model named in the Hugging Face card metadata (base_model_relation finetune).

Other releases in the gpt-oss-safeguard family

gpt-oss-safeguard family overview

What this catalog does not know

Unknown means the sources reviewed for this record do not document it. It is not evidence that something does not exist.
  • Training code: unknown.
  • Training-data information: unknown.
  • Training-data access: unknown.
  • Complete training pipeline: unknown.
  • Legacy data assessment (v0.1): unknown.
  • Training recipe: unknown.

Have a primary source? How to report a correction.

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

Developed and published by OpenAI. Its operating company, OpenAI Group PBC, is described as a Delaware public benefit corporation at 1455 3rd Street, San Francisco, California, in a February 2026 agreement filed with the SEC.610

Assessed Oct 1, 2026

Sources

  1. 1.
    openai/gpt-oss-safeguard-120b model card (external site: huggingface.co)

    OpenAI (Hugging Face) · Model card · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  2. 2.
    Hugging Face model metadata for openai/gpt-oss-safeguard-120b (external site: huggingface.co)

    Hugging Face · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  3. 3.
    openai/gpt-oss-safeguard-120b config.json (external site: huggingface.co)

    OpenAI (Hugging Face) · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  4. 4.
    openai/gpt-oss-safeguard-120b LICENSE (external site: huggingface.co)

    OpenAI (Hugging Face) · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  5. 5.
    openai/gpt-oss-safeguard-120b USAGE_POLICY (external site: huggingface.co)

    OpenAI (Hugging Face) · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  6. 6.
    Technical Report: Performance and baseline evaluations of gpt-oss-safeguard-120b and gpt-oss-safeguard-20b (external site: cdn.openai.com)

    OpenAI · Paper · published Oct 29, 2025 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  7. 7.
    openai/gpt-oss-safeguard README (external site: raw.githubusercontent.com)

    OpenAI · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  8. 8.
    openai/gpt-oss-safeguard LICENSE (external site: raw.githubusercontent.com)

    OpenAI · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  9. 9.
    gpt-oss-safeguard user guide (OpenAI Cookbook) (external site: developers.openai.com)

    OpenAI · Documentation · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  10. 10.
    Exhibit 10.1: Equity commitment letter agreement between OpenAI Group PBC and Amazon (external site: sec.gov)

    U.S. Securities and Exchange Commission (Amazon.com, Inc. filing) · Filing · published Feb 27, 2026 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project