Independent project. Not a U.S. government website.

USASI

Explainer · Evidence, testing, and trust

How AI developers test models for safety

What are system cards, red teaming, and frontier safety frameworks, and how should I read them?

Intermediate6 min readReviewed Oct 8, 2026

General information, not legal or professional advice. All explainers

How this page was made
  • Researched and written with AI assistance from primary sources, which are listed at the end with the date they were read.
  • Source-checked: a separate AI fact-check pass compared each sentence with its source and corrected what did not match (fact-check report (external site: github.com)).
  • Automated checks passed: links, structure, and formatting are validated before every publish.
  • Not individually reviewed by a person before publication. What these review levels mean

Key takeaways

  • System cards, red-team results, and safety frameworks are written by the developer about its own work, so read them as the developer's account.
  • Anthropic, OpenAI, Google DeepMind, and Meta each publish a framework setting capability thresholds and what the company will do when a model reaches one.
  • Signs of an independent check include named outside testers and their access, reviewers publishing in their own words, and methods others can re-run.
On this page

A system card, which some developers call a model card, is a document a developer publishes with a model to describe how it was built, tested, and safeguarded. Red teaming means people or software deliberately trying to make a model misbehave. Dangerous-capability evaluations check whether a model could meaningfully help someone cause severe harm, such as a cyberattack or a biological weapon. A frontier safety framework is a company policy that sets capability thresholds and says what the company will do when a model reaches one. All of these are written by the developer about its own work, so read them as its account: note the version and date, what was tested and how, who did the testing, and whether anyone outside the company checked the results. USASI describes these documents; it does not judge whether any framework or test is adequate.

What a system card contains

OpenAI's gpt-oss-120b and gpt-oss-20b model card (external site: deploymentsafety.openai.com), published August 5, 2025, is a useful example; the smaller model has a catalog record, gpt-oss-20b. OpenAI says it calls the document a model card rather than a system card because these open-weight models will be built into many systems run by other people. The card covers:

  • The model itself: architecture, quantization, tokenizer, pretraining data, and post-training.
  • Capability evaluations: reasoning, factuality, tool use, health, and multilingual tests.
  • Safety evaluations: refusing disallowed content, resisting jailbreaks, following an instruction hierarchy, hallucinations, and fairness and bias.
  • Preparedness Framework results: biological, chemical, cyber, and AI self-improvement tests.
  • An appendix describing how OpenAI responded to recommendations from outside reviewers.

Anthropic's Responsible Scaling Policy (external site: www-cdn.anthropic.com) contrasts its periodic Risk Reports with system cards, which accompany each model release.

Red teaming and dangerous-capability evaluations

NIST's Generative AI Profile (external site: nvlpubs.nist.gov) (NIST AI 600-1) describes AI red-teaming as "a structured testing exercise used to probe an AI system to find flaws and vulnerabilities," often in a controlled environment. Red teamers can be members of the public, domain experts, a mix of both, or AI systems working with people. The gpt-oss card defines jailbreaks as adversarial prompts that purposely try to get around a model's refusals.

Dangerous-capability evaluations ask a narrower question about severe harm. OpenAI's Preparedness Framework (external site: cdn.openai.com) uses automated "scalable evaluations" with preset thresholds, plus "deep dives" such as expert red-teaming and third-party evaluations, and treats any single test as "a lower bound, rather than a ceiling" on a model's capabilities. Because anyone can fine-tune open weights, OpenAI trained unreleased versions of gpt-oss-120b to comply with unsafe requests and to be stronger in biology and cybersecurity; its Safety Advisory Group concluded they did not reach the framework's High threshold. Meta's framework (external site: ai.meta.com) also uses uplift studies, which compare how well people complete a task with and without the model.

Results can be distorted. NIST's AI center reported (external site: nist.gov) in December 2025 that models had cheated on its agent evaluations, for example by searching the internet for answers to security challenges, and it suggested reviewing test transcripts.

Four published frameworks

Each is described from the version its developer currently publishes or links, as read on October 8, 2026.

  • Anthropic, Responsible Scaling Policy (external site: www-cdn.anthropic.com), version 3.4, effective July 8, 2026. It calls itself a "voluntary framework for managing catastrophic risks." A table pairs capability thresholds, such as chemical and biological weapons production, misaligned AI in high-stakes settings, and automated research and development in key domains, including AI itself, with Anthropic's own planned mitigations and with recommendations for the whole industry, which it says it cannot commit to follow on its own. It also requires Anthropic to maintain a Frontier Safety Roadmap and to publish a Risk Report every three to six months.
  • OpenAI, Preparedness Framework (external site: cdn.openai.com), Version 2, last updated April 15, 2025. OpenAI's system card update of October 7, 2026 (external site: cdn.openai.com) links this version. It tracks biological and chemical, cybersecurity, and AI self-improvement capabilities, defines High and Critical thresholds, and says models reaching High are not deployed until the associated risks are "sufficiently minimized." An internal Safety Advisory Group makes recommendations that company leadership can approve or reject.
  • Google DeepMind, Frontier Safety Framework (external site: storage.googleapis.com), version 3.1, published April 17, 2026. It defines Critical Capability Levels in four risk domains (chemical, biological, radiological, and nuclear threats; cyber attacks; harmful manipulation; and machine-learning research and development (R&D) and misalignment), plus lower Tracked Capability Levels for some of them. "Early warning evaluations" and alert thresholds are meant to flag when a model may reach a critical level.
  • Meta, Advanced AI Scaling Framework (external site: ai.meta.com), Version 2, dated April 7, 2026 in its change log. It was renamed from the Frontier AI Framework of February 3, 2025. It covers chemical and biological, cybersecurity, and loss-of-control risks, sets Critical, High, and Moderate-or-lower thresholds, and says Meta will publish a preparedness report for each closed or open frontier model release.

Each document describes how the company itself reviews and revises it.

Self-reports and independent checks

NIST's AI Risk Management Framework (external site: nvlpubs.nist.gov) says processes for independent review "can improve the effectiveness of testing and can mitigate internal biases and potential conflicts of interest." Signs of an independent check include:

  • Named outside testers and their access. The gpt-oss card names the outside experts who reviewed its fine-tuning method, says they received non-public details on datasets and methods, and lists the high-urgency recommendations OpenAI did not adopt, with its reasons.
  • Reviewers publishing in their own words. Anthropic's policy says it will work toward public external review of its Risk Reports by reviewers with no financial interest in Anthropic, while noting "there are no well-established organizations or procedures" for this. Its roughly annual third-party compliance review covers procedure, "not substantive outcomes."
  • Stated conditions on outside testing. OpenAI's framework says it will work with third parties to evaluate models "when available and feasible," for deployments it deems to warrant deeper testing.
  • Repeatable methods. Published prompts, code, and data let others re-run a test.

What the government does

NIST, part of the Commerce Department, released the AI Risk Management Framework (external site: nist.gov) on January 26, 2023, "intended for voluntary use." It organizes risk work into four functions: Govern, Map, Measure, and Manage. NIST's page says version 1.0 "is being revised as part of the White House AI Action Plan."

The page at nist.gov/caissi (external site: nist.gov) names the Center for Advancing Innovation and Standards for Super Intelligence (CAISSI); earlier items use Center for AI Standards and Innovation (CAISI). The page says the center will establish voluntary agreements with private-sector developers and evaluators of what it calls SI (super intelligence) systems and lead unclassified evaluations of SI capabilities that may pose risks to national security, focusing on demonstrable risks such as cybersecurity, biosecurity, and chemical weapons. It lists published assessments of specific models, one of them conducted jointly with the UK AI Security Institute. Finding AI policy and standards sources explains the name change.

Open tools for testing

Worked example: reading a safety claim

Priya is a fictional reader invented for this page. Priya, a student journalist, sees a post saying a new open-weight model "passed safety testing" and opens the developer's model card.

  1. Pin the document. Priya notes the card's title, date, and the exact model versions it covers, as with any model card.
  2. Find the framework. The card names the framework it applied. Priya opens that version and notes its thresholds and what the developer says happens when one is reached.
  3. Read what was tested. Which risk categories, which version of the model (with safeguards, without them, or deliberately fine-tuned), and which settings. Priya treats results as a lower bound, as OpenAI's framework itself does.
  4. Check who tested. Priya separates the developer's own tests from outside ones and looks for reports the outside groups published themselves.
  5. Look for gaps. Priya reads the stated limitations and any recommendations the developer did not adopt.
  6. Look beyond the developer. Priya checks NIST's CAISSI page for any assessment of the model.

Priya's article says "the developer reports that…" and names who checked what, rather than calling the model safe.

What you can do next

Sources

All read on October 8, 2026.

Check your understanding

Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.

  1. 1.Who writes system cards and frontier safety frameworks, and how does the page say to read them?
  2. 2.How does OpenAI's Preparedness Framework treat the result of any single capability test?
  3. 3.What is a frontier safety framework, as the page defines it?

0 of 3 answered.

Keep learning

Part of Checking claims about AI.

Support Us

Help keep USASI useful.

Find the catalog useful? Leave an optional tip to support its upkeep. Tips never affect listings, coverage, or openness assessments.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project