Llama Guard 4 12B
Release in the Llama family · version Llama-Guard-4-12B
Llama Guard 4 is a 12-billion-parameter safety classifier from Meta for text and images. Given a prompt sent to a language model, or a model's response, it generates text saying whether the content is safe or unsafe and, if unsafe, which categories it violates: 13 categories based on the MLCommons hazards taxonomy plus a text-only category for code interpreter abuse. It is a dense model pruned from the pre-trained Llama 4 Scout and fine-tuned for classification, and it accepts multiple images in one prompt. Meta released it on April 29, 2025, as one of its Llama protection tools.18
- Documentation: Model card (PurpleLlama repository) (external site: github.com)
- Model hub: Hugging Face (gated) (external site: huggingface.co)
- Documentation: Prompt format documentation (external site: dev.meta.ai)
- License: Llama 4 Community License Agreement (external site: github.com)
- Release notes: Release announcement (external site: ai.meta.com)
Availability and license
Overall availability
The Hugging Face repository is gated with manual approval: requesters must accept the Llama 4 Community License and submit their name, date of birth, country, affiliation, and job title. The Llama 4 Acceptable Use Policy shipped with the model withholds the license grant for Llama 4 multimodal models from individuals domiciled in, and companies with a principal place of business in, the European Union; this catalog has not confirmed how Meta applies that clause to Llama Guard 4.453
Availability is separate from permission: read the license before using or redistributing.
Llama 4 Community License Agreement (external site: github.com)24
Plain-language guide to the Llama 4 Community License Agreement
Custom Meta license, effective April 5, 2025; the Hugging Face metadata lists it as "other" with the name llama4. It grants a royalty-free, non-exclusive license to use, modify, and redistribute. Redistributors must include the agreement, show "Built with Llama", and keep an attribution notice, and distributed models trained or fine-tuned with Llama materials or outputs must start their names with "Llama". Use must follow the incorporated Llama 4 Acceptable Use Policy. Licensees whose products had more than 700 million monthly active users in the month before the Llama 4 release date must request a separate license. The license ends for anyone who sues alleging that the Llama materials or outputs infringe their rights. California law governs.243
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Unknown — no complete fact-level rights review
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Model-disclosure tier
The weights can be obtained only by request, with approval, or by some users — for example a gated download that the publisher reviews. Not counted as open-weight.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Partial | BF16 safetensors (12,001,097,216 parameters per the repository metadata) in a Hugging Face repository gated with manual approval and a personal-information form.45 |
| Inference codeIs code for running the model published? | Public | The repository metadata identifies the Hugging Face Transformers Llama 4 architecture (Llama4ForConditionalGeneration), and Meta's documentation specifies the prompt format, including how images are split into 336-by-336-pixel tiles. This review could not read the gated Hugging Face card.46 |
| Training codeIs the code used to train the model published? | Unknown | No training or pruning code for Llama Guard 4 was found in the repositories reviewed. |
| Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness. | Partial | The card says post-training used the Llama Guard 3 8B and 11B-vision training data plus multi-image samples (mostly two to five images) and multilingual data, written by human annotators or translated from English, at roughly three parts text-only to one part multimodal. The data is not released.1 |
| Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable. | Unknown | Not assessed. |
| Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient. | Unknown | Not assessed. |
| Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above. | Unknown | Not assessed. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The card describes pruning the Llama 4 Scout pre-trained checkpoint by removing all routed experts and routers and keeping the shared expert, with no further pre-training, followed by post-training for classification. Hyperparameters are not given.1 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The card reports recall, false positive rate, and F1 against Llama Guard 3 on an in-house test set that the card does not make available.1 |
What it is useful for
Filtering the inputs to a generative model, its outputs, or both, as part of a larger system. Meta's documentation says it is designed to work with Llama 4 Scout and Maverick and can be used as a drop-in replacement for Llama Guard 3 8B and 11B. Meta's Llama Protections page says it is also available through the moderations endpoint of Meta's Llama API.167
Run and use notes
- Meta's card and documentation say the model can be run on a single GPU; neither gives a memory figure or a precision for that statement.16
- The card says the model was tested mostly with prompts containing a few images (most often three), that categories such as defamation, intellectual property, and elections may need up-to-date facts to judge, and that as a language model it may be open to prompt injection; it points to Meta's Prompt Guard 2 for detecting prompt attacks.1
Organization context
Provenance and derivatives
Meta built Llama Guard 4 by pruning the pre-trained Llama 4 Scout mixture-of-experts checkpoint into a dense model and post-training it for safety classification. It shares Llama 4 Scout and Maverick's tokenizer and vision encoder.1
- Derived from: Llama 4 Scout (pre-trained) — Routed experts and routers removed; the shared expert became a dense feed-forward layer.
Other releases in the Llama family
- Llama 3.1 8BModel-disclosure tier (USASI rubric v0.2): Restricted weights
- Llama 4 Maverick (17Bx128E)Model-disclosure tier (USASI rubric v0.2): Restricted weights
- Llama 4 Scout (17Bx16E)Model-disclosure tier (USASI rubric v0.2): Restricted weights
What this catalog does not know
- Training code: unknown.
- Training-data access: unknown.
- Complete training pipeline: unknown.
- Legacy data assessment (v0.1): unknown.
Have a primary source? How to report a correction.
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.