Key takeaways
- System cards, red-team results, and safety frameworks are written by the developer about its own work, so read them as the developer's account.
- Anthropic, OpenAI, Google DeepMind, and Meta each publish a framework setting capability thresholds and what the company will do when a model reaches one.
- Signs of an independent check include named outside testers and their access, reviewers publishing in their own words, and methods others can re-run.
On this page
A system card, which some developers call a model card, is a document a developer publishes with a model to describe how it was built, tested, and safeguarded. Red teaming means people or software deliberately trying to make a model misbehave. Dangerous-capability evaluations check whether a model could meaningfully help someone cause severe harm, such as a cyberattack or a biological weapon. A frontier safety framework is a company policy that sets capability thresholds and says what the company will do when a model reaches one. All of these are written by the developer about its own work, so read them as its account: note the version and date, what was tested and how, who did the testing, and whether anyone outside the company checked the results. USASI describes these documents; it does not judge whether any framework or test is adequate.
What a system card contains
OpenAI's gpt-oss-120b and gpt-oss-20b model card (external site: deploymentsafety.openai.com), published August 5, 2025, is a useful example; the smaller model has a catalog record, gpt-oss-20b. OpenAI says it calls the document a model card rather than a system card because these open-weight models will be built into many systems run by other people. The card covers:
- The model itself: architecture, quantization, tokenizer, pretraining data, and post-training.
- Capability evaluations: reasoning, factuality, tool use, health, and multilingual tests.
- Safety evaluations: refusing disallowed content, resisting jailbreaks, following an instruction hierarchy, hallucinations, and fairness and bias.
- Preparedness Framework results: biological, chemical, cyber, and AI self-improvement tests.
- An appendix describing how OpenAI responded to recommendations from outside reviewers.
Anthropic's Responsible Scaling Policy (external site: www-cdn.anthropic.com) contrasts its periodic Risk Reports with system cards, which accompany each model release.
Red teaming and dangerous-capability evaluations
NIST's Generative AI Profile (external site: nvlpubs.nist.gov) (NIST AI 600-1) describes AI red-teaming as "a structured testing exercise used to probe an AI system to find flaws and vulnerabilities," often in a controlled environment. Red teamers can be members of the public, domain experts, a mix of both, or AI systems working with people. The gpt-oss card defines jailbreaks as adversarial prompts that purposely try to get around a model's refusals.
Dangerous-capability evaluations ask a narrower question about severe harm. OpenAI's Preparedness Framework (external site: cdn.openai.com) uses automated "scalable evaluations" with preset thresholds, plus "deep dives" such as expert red-teaming and third-party evaluations, and treats any single test as "a lower bound, rather than a ceiling" on a model's capabilities. Because anyone can fine-tune open weights, OpenAI trained unreleased versions of gpt-oss-120b to comply with unsafe requests and to be stronger in biology and cybersecurity; its Safety Advisory Group concluded they did not reach the framework's High threshold. Meta's framework (external site: ai.meta.com) also uses uplift studies, which compare how well people complete a task with and without the model.
Results can be distorted. NIST's AI center reported (external site: nist.gov) in December 2025 that models had cheated on its agent evaluations, for example by searching the internet for answers to security challenges, and it suggested reviewing test transcripts.
Four published frameworks
Each is described from the version its developer currently publishes or links, as read on October 8, 2026.
- Anthropic, Responsible Scaling Policy (external site: www-cdn.anthropic.com), version 3.4, effective July 8, 2026. It calls itself a "voluntary framework for managing catastrophic risks." A table pairs capability thresholds, such as chemical and biological weapons production, misaligned AI in high-stakes settings, and automated research and development in key domains, including AI itself, with Anthropic's own planned mitigations and with recommendations for the whole industry, which it says it cannot commit to follow on its own. It also requires Anthropic to maintain a Frontier Safety Roadmap and to publish a Risk Report every three to six months.
- OpenAI, Preparedness Framework (external site: cdn.openai.com), Version 2, last updated April 15, 2025. OpenAI's system card update of October 7, 2026 (external site: cdn.openai.com) links this version. It tracks biological and chemical, cybersecurity, and AI self-improvement capabilities, defines High and Critical thresholds, and says models reaching High are not deployed until the associated risks are "sufficiently minimized." An internal Safety Advisory Group makes recommendations that company leadership can approve or reject.
- Google DeepMind, Frontier Safety Framework (external site: storage.googleapis.com), version 3.1, published April 17, 2026. It defines Critical Capability Levels in four risk domains (chemical, biological, radiological, and nuclear threats; cyber attacks; harmful manipulation; and machine-learning research and development (R&D) and misalignment), plus lower Tracked Capability Levels for some of them. "Early warning evaluations" and alert thresholds are meant to flag when a model may reach a critical level.
- Meta, Advanced AI Scaling Framework (external site: ai.meta.com), Version 2, dated April 7, 2026 in its change log. It was renamed from the Frontier AI Framework of February 3, 2025. It covers chemical and biological, cybersecurity, and loss-of-control risks, sets Critical, High, and Moderate-or-lower thresholds, and says Meta will publish a preparedness report for each closed or open frontier model release.
Each document describes how the company itself reviews and revises it.
Self-reports and independent checks
NIST's AI Risk Management Framework (external site: nvlpubs.nist.gov) says processes for independent review "can improve the effectiveness of testing and can mitigate internal biases and potential conflicts of interest." Signs of an independent check include:
- Named outside testers and their access. The gpt-oss card names the outside experts who reviewed its fine-tuning method, says they received non-public details on datasets and methods, and lists the high-urgency recommendations OpenAI did not adopt, with its reasons.
- Reviewers publishing in their own words. Anthropic's policy says it will work toward public external review of its Risk Reports by reviewers with no financial interest in Anthropic, while noting "there are no well-established organizations or procedures" for this. Its roughly annual third-party compliance review covers procedure, "not substantive outcomes."
- Stated conditions on outside testing. OpenAI's framework says it will work with third parties to evaluate models "when available and feasible," for deployments it deems to warrant deeper testing.
- Repeatable methods. Published prompts, code, and data let others re-run a test.
What the government does
NIST, part of the Commerce Department, released the AI Risk Management Framework (external site: nist.gov) on January 26, 2023, "intended for voluntary use." It organizes risk work into four functions: Govern, Map, Measure, and Manage. NIST's page says version 1.0 "is being revised as part of the White House AI Action Plan."
The page at nist.gov/caissi (external site: nist.gov) names the Center for Advancing Innovation and Standards for Super Intelligence (CAISSI); earlier items use Center for AI Standards and Innovation (CAISI). The page says the center will establish voluntary agreements with private-sector developers and evaluators of what it calls SI (super intelligence) systems and lead unclassified evaluations of SI capabilities that may pose risks to national security, focusing on demonstrable risks such as cybersecurity, biosecurity, and chemical weapons. It lists published assessments of specific models, one of them conducted jointly with the UK AI Security Institute. Finding AI policy and standards sources explains the name change.
Open tools for testing
- Dioptra is NIST's open-source (external site: pages.nist.gov) test platform. Its README (external site: github.com) says it supports the AI Risk Management Framework's Measure function and lists uses including third-party audits and red-teaming in a controlled environment.
- Petri runs automated audits: an auditor model holds multi-turn conversations with the model under test, starting from seed instructions, and a judge model scores the transcripts. Its documentation (external site: meridianlabs-ai.github.io) describes it as a collaboration between Meridian Labs and the UK AISI Red Team, based on work by Anthropic.
- gpt-oss-safeguard is a pair of open-weight models that classify text against safety policies you supply; OpenAI's README (external site: github.com) says they are intended for safety use cases.
Worked example: reading a safety claim
Priya is a fictional reader invented for this page. Priya, a student journalist, sees a post saying a new open-weight model "passed safety testing" and opens the developer's model card.
- Pin the document. Priya notes the card's title, date, and the exact model versions it covers, as with any model card.
- Find the framework. The card names the framework it applied. Priya opens that version and notes its thresholds and what the developer says happens when one is reached.
- Read what was tested. Which risk categories, which version of the model (with safeguards, without them, or deliberately fine-tuned), and which settings. Priya treats results as a lower bound, as OpenAI's framework itself does.
- Check who tested. Priya separates the developer's own tests from outside ones and looks for reports the outside groups published themselves.
- Look for gaps. Priya reads the stated limitations and any recommendations the developer did not adopt.
- Look beyond the developer. Priya checks NIST's CAISSI page for any assessment of the model.
Priya's article says "the developer reports that…" and names who checked what, rather than calling the model safe.
What you can do next
- Browse the AI safety, security, and trust hub and the evaluation hub, and read how to read an AI evaluation.
- Open the records for Dioptra, Petri, and gpt-oss-safeguard, and the NIST profile.
- Look up model card, benchmark, evaluation harness, and fine-tuning in the glossary.
- Read labels, watermarks, and content credentials to see how developers mark what their models generate.
Sources
All read on October 8, 2026.
- OpenAI: gpt-oss-120b and gpt-oss-20b Model Card (external site: deploymentsafety.openai.com) (published August 5, 2025); Preparedness Framework, Version 2 (external site: cdn.openai.com) (last updated April 15, 2025); GPT-6 Sol and GPT-6 Luna: October 2026 update (external site: cdn.openai.com) (dated October 7, 2026); gpt-oss-safeguard README (external site: github.com)
- Anthropic: Responsible Scaling Policy page (external site: anthropic.com) (last updated August 14, 2026) and Responsible Scaling Policy, version 3.4 (external site: www-cdn.anthropic.com) (effective July 8, 2026)
- Google DeepMind: Frontier safety page (external site: deepmind.google) and Frontier Safety Framework, version 3.1 (external site: storage.googleapis.com) (published April 17, 2026)
- Meta: Advanced AI Scaling Framework, Version 2 (external site: ai.meta.com)
- NIST: AI Risk Management Framework page (external site: nist.gov); AI RMF 1.0 (NIST AI 100-1) (external site: nvlpubs.nist.gov); Generative AI Profile (NIST AI 600-1) (external site: nvlpubs.nist.gov); CAISSI (external site: nist.gov); Cheating On AI Agent Evaluations (external site: nist.gov) (December 2, 2025); Dioptra README (external site: github.com) and documentation (external site: pages.nist.gov)
- Meridian Labs: Inspect Petri documentation (external site: meridianlabs-ai.github.io)
Check your understanding
Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.
0 of 3 answered.