Topic hub
AI safety, security, and trust
How developers, standards bodies, and independent groups test AI systems for harm and misuse: safety frameworks and model cards, testing tools, safety classifiers, content provenance and watermarking, and prompt injection.
NIST's AI Risk Management Framework, released on January 26, 2023, is intended for voluntary use, and its core has four functions: Govern, Map, Measure, and Manage. NIST says the framework is being revised as part of the White House AI Action Plan. NIST's Dioptra test platform supports the Measure function and lists red-teaming in a controlled environment among its uses. NIST's Center for Advancing Innovation and Standards for Super Intelligence (CAISSI) says it will help industry develop voluntary standards and lead unclassified evaluations focused on demonstrable risks such as cybersecurity, biosecurity, and chemical weapons.1234
Several developers publish frameworks for testing new models for risks of severe harm. OpenAI's Preparedness Framework (version 2, April 2025) tracks biological and chemical, cybersecurity, and AI self-improvement capabilities against set thresholds. Google DeepMind's Frontier Safety Framework uses Critical Capability Levels, thresholds at which a model may pose heightened risk of severe harm without mitigations. Anthropic's Responsible Scaling Policy is at version 3.4, effective July 8, 2026. Model cards report how one release was tested: OpenAI's gpt-oss card says that because open weights can be fine-tuned, OpenAI also tested adversarially fine-tuned versions of gpt-oss-120b, which it did not release.5678
Testing tools probe models for unwanted behavior. In Petri, an auditor model runs multi-turn scenarios with a target model and a judge model scores the transcripts. Anthropic says Petri has been part of its alignment assessment of every Claude model since Claude Sonnet 4.5, and in May 2026 it handed Petri's development to Meridian Labs, an AI evaluation nonprofit. METR's Inspect-Hawk runs evaluations written with Inspect AI, the UK AI Security Institute's open-source framework, in isolated Kubernetes pods. MLCommons' AILuminate benchmarks cover 12 hazard categories and include jailbreak tests with text and with text plus images.9101112
Safety classifiers screen what goes into and comes out of another model. Meta's Llama Guard 4 is a 12-billion-parameter classifier for text and images that labels a prompt or a response as safe or unsafe and, if unsafe, lists the violated categories from a taxonomy based on MLCommons' hazard categories. OpenAI's gpt-oss-safeguard models, built on gpt-oss, instead classify text against a written policy that the developer supplies. OpenAI says they are intended for safety use cases and work correctly only with its harmony response format.131415
Provenance records where content came from. NIST's report on synthetic content (NIST AI 100-4) covers watermarks, metadata, and detection, and says none of these techniques is a comprehensive solution on its own. The C2PA standard's Content Credentials are a cryptographically bound record of an asset's origin, edits, and use of AI; they show whether that record is intact and signed by a trusted party, not whether its claims are true. Google DeepMind has published watermarking code for generated text (SynthID Text, a reference implementation not meant for production) and for AI-generated biological sequences and structures (SynthID Bio).16171819
Prompt injection is a security risk for systems built on language models. NIST's taxonomy of attacks (NIST AI 100-2 E2025) calls it direct when the system's main user supplies instructions that are appended to higher-trust ones, such as the system prompt, and indirect when a third party plants instructions in data the system reads, such as content fetched by an agent or a retrieval system. NIST says current mitigations do not offer full protection, so designers may assume injection is possible whenever a model sees untrusted input. For agents that use tools, the added risks include running arbitrary code and exfiltrating data.20
What this hub covers
Covers how AI systems are tested and governed for safety and security: government frameworks and standards work, developer safety frameworks and model cards, testing and red-teaming tools, safety classifiers, content provenance and watermarking, and attacks such as prompt injection. It does not rate how safe any model is or report test scores. Featured records are examples chosen to cover the subject, not a ranking or a complete list.
Reading path
- Safety testing and frameworksStart here for how developers and outside groups test models before and after release.
- How to read a model cardModel and system cards are where a release's safety testing is usually reported.
- Reading evaluationsQuestions to ask of any published test result, including safety benchmarks.
- AI content provenanceHow watermarks, metadata, and Content Credentials record where media came from.
- How models use toolsTool use is where prompt injection can turn into actions, so read this before trusting an agent.
- Policy and standards sourcesHow to find and read government frameworks and standards documents.
- Acceptable use policyA publisher's list of prohibited uses, which a license can make binding by reference.
- Benchmarks and evaluationEvaluation harnesses and benchmarks, including safety and system-performance tests.
- AI agents and agent toolingAgent protocols and their consent and permission rules.
In the catalog
- Records that mention safetyA text search that matches safety classifiers, alignment audits, and safety benchmarks.
- Records that mention securityA text search that matches security testing tools and benchmarks.
- Evaluation tools and benchmarksOpen evaluation records, including safety and jailbreak benchmarks.
- Standards bodiesOrganizations whose records list standards work, including NIST and MLCommons.
- Nonprofit research organizationsNonprofits that do research, several of them on AI safety and evaluation.
Featured records
Examples chosen to cover the subject; not a ranking or a complete list.
- National Institute of Standards and TechnologyOrganization
- METR (Model Evaluation & Threat Research)Organization
- MLCommonsOrganization
- Center for AI Safety (CAIS)Organization
- FAR.AIOrganization
- DioptraOpen model or tool
- Petri (Inspect Petri)Open model or tool
- Inspect-Hawk (Hawk)Open model or tool
- AILuminateOpen model or tool
- gpt-oss-safeguardOpen model or tool
- gpt-ossOpen model or tool
- SynthID BioOpen model or tool
- Llama Guard 4 12BOpen model or tool
- SynthID TextOpen model or tool
Primary documents
- NIST AI 100-1: Artificial Intelligence Risk Management Framework (AI RMF 1.0) (external site: nvlpubs.nist.gov)National Institute of Standards and TechnologyThe AI Risk Management Framework itself, including the four core functions and the characteristics NIST uses to describe trustworthy AI.
- NIST AI 100-2 E2025: Adversarial Machine Learning, A Taxonomy and Terminology of Attacks and Mitigations (external site: nvlpubs.nist.gov)National Institute of Standards and TechnologyNIST's taxonomy of attacks on machine learning systems, with definitions of direct and indirect prompt injection and a statement of the limits of current mitigations.
- NIST AI 100-4: Reducing Risks Posed by Synthetic Content (external site: nvlpubs.nist.gov)National Institute of Standards and TechnologyAn overview of watermarking, metadata, and detection for synthetic content, with the trade-offs and research gaps of each approach.
- C2PA and Content Credentials Explainer (version 2.4) (external site: spec.c2pa.org)Coalition for Content Provenance and Authenticity (C2PA)The C2PA's plain-language explainer for Content Credentials, including what a valid credential does and does not tell you.
- Preparedness Framework, version 2 (external site: cdn.openai.com)OpenAIOne developer's full framework, with its tracked categories, capability thresholds, and safeguard requirements, useful for reading model cards that cite it.
- gpt-oss-120b and gpt-oss-20b Model Card (external site: cdn.openai.com)OpenAIA model card for an open-weight release that explains how the developer tested adversarially fine-tuned versions, a step specific to models whose weights are public.
- Anthropic's Responsible Scaling Policy (external site: anthropic.com)AnthropicLists every version of Anthropic's Responsible Scaling Policy with redlines and change notes, which shows how such a policy changes over time.
- Llama Guard 4 model card (external site: raw.githubusercontent.com)Meta, PurpleLlama (GitHub)Shows the hazard categories a safety classifier is trained on and how it reports which category a prompt or response violates.
- Center for Advancing Innovation and Standards for Super Intelligence (CAISSI) (external site: nist.gov)National Institute of Standards and TechnologyNIST's page for CAISSI, with its stated responsibilities and links to its published evaluations of specific models. The evaluation items listed there use the name Center for AI Standards and Innovation (CAISI).
- Dioptra README (external site: raw.githubusercontent.com)NIST (GitHub)NIST's test platform for AI, with its intended uses from in-house testing to audits and red-teaming.
Sources · reviewed Oct 8, 2026
Support Us
Help keep USASI useful.
Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).
About supporting this project