Independent project. Not a U.S. government website.

USASI

Topic hub

AI safety, security, and trust

How developers, standards bodies, and independent groups test AI systems for harm and misuse: safety frameworks and model cards, testing tools, safety classifiers, content provenance and watermarking, and prompt injection.

NIST's AI Risk Management Framework, released on January 26, 2023, is intended for voluntary use, and its core has four functions: Govern, Map, Measure, and Manage. NIST says the framework is being revised as part of the White House AI Action Plan. NIST's Dioptra test platform supports the Measure function and lists red-teaming in a controlled environment among its uses. NIST's Center for Advancing Innovation and Standards for Super Intelligence (CAISSI) says it will help industry develop voluntary standards and lead unclassified evaluations focused on demonstrable risks such as cybersecurity, biosecurity, and chemical weapons.1234

Several developers publish frameworks for testing new models for risks of severe harm. OpenAI's Preparedness Framework (version 2, April 2025) tracks biological and chemical, cybersecurity, and AI self-improvement capabilities against set thresholds. Google DeepMind's Frontier Safety Framework uses Critical Capability Levels, thresholds at which a model may pose heightened risk of severe harm without mitigations. Anthropic's Responsible Scaling Policy is at version 3.4, effective July 8, 2026. Model cards report how one release was tested: OpenAI's gpt-oss card says that because open weights can be fine-tuned, OpenAI also tested adversarially fine-tuned versions of gpt-oss-120b, which it did not release.5678

Testing tools probe models for unwanted behavior. In Petri, an auditor model runs multi-turn scenarios with a target model and a judge model scores the transcripts. Anthropic says Petri has been part of its alignment assessment of every Claude model since Claude Sonnet 4.5, and in May 2026 it handed Petri's development to Meridian Labs, an AI evaluation nonprofit. METR's Inspect-Hawk runs evaluations written with Inspect AI, the UK AI Security Institute's open-source framework, in isolated Kubernetes pods. MLCommons' AILuminate benchmarks cover 12 hazard categories and include jailbreak tests with text and with text plus images.9101112

Safety classifiers screen what goes into and comes out of another model. Meta's Llama Guard 4 is a 12-billion-parameter classifier for text and images that labels a prompt or a response as safe or unsafe and, if unsafe, lists the violated categories from a taxonomy based on MLCommons' hazard categories. OpenAI's gpt-oss-safeguard models, built on gpt-oss, instead classify text against a written policy that the developer supplies. OpenAI says they are intended for safety use cases and work correctly only with its harmony response format.131415

Provenance records where content came from. NIST's report on synthetic content (NIST AI 100-4) covers watermarks, metadata, and detection, and says none of these techniques is a comprehensive solution on its own. The C2PA standard's Content Credentials are a cryptographically bound record of an asset's origin, edits, and use of AI; they show whether that record is intact and signed by a trusted party, not whether its claims are true. Google DeepMind has published watermarking code for generated text (SynthID Text, a reference implementation not meant for production) and for AI-generated biological sequences and structures (SynthID Bio).16171819

Prompt injection is a security risk for systems built on language models. NIST's taxonomy of attacks (NIST AI 100-2 E2025) calls it direct when the system's main user supplies instructions that are appended to higher-trust ones, such as the system prompt, and indirect when a third party plants instructions in data the system reads, such as content fetched by an agent or a retrieval system. NIST says current mitigations do not offer full protection, so designers may assume injection is possible whenever a model sees untrusted input. For agents that use tools, the added risks include running arbitrary code and exfiltrating data.20

What this hub covers

Covers how AI systems are tested and governed for safety and security: government frameworks and standards work, developer safety frameworks and model cards, testing and red-teaming tools, safety classifiers, content provenance and watermarking, and attacks such as prompt injection. It does not rate how safe any model is or report test scores. Featured records are examples chosen to cover the subject, not a ranking or a complete list.

Reading path

  1. 1.Safety testing and frameworksStart here for how developers and outside groups test models before and after release.
  2. 2.How to read a model cardModel and system cards are where a release's safety testing is usually reported.
  3. 3.Reading evaluationsQuestions to ask of any published test result, including safety benchmarks.
  4. 4.AI content provenanceHow watermarks, metadata, and Content Credentials record where media came from.
  5. 5.How models use toolsTool use is where prompt injection can turn into actions, so read this before trusting an agent.
  6. 6.Policy and standards sourcesHow to find and read government frameworks and standards documents.
  7. 7.Acceptable use policyA publisher's list of prohibited uses, which a license can make binding by reference.
  8. 8.Benchmarks and evaluationEvaluation harnesses and benchmarks, including safety and system-performance tests.
  9. 9.AI agents and agent toolingAgent protocols and their consent and permission rules.

In the catalog

Examples chosen to cover the subject; not a ranking or a complete list.

Primary documents

Sources · reviewed Oct 8, 2026

  1. 1.
    AI Risk Management Framework (external site: nist.gov)

    National Institute of Standards and Technology · Official page · accessed Oct 8, 2026

  2. 2.
    NIST AI 100-1: Artificial Intelligence Risk Management Framework (AI RMF 1.0) (external site: nvlpubs.nist.gov)

    National Institute of Standards and Technology · Official page · published Jan 2023 · accessed Oct 8, 2026

  3. 3.
    Dioptra README (external site: raw.githubusercontent.com)

    NIST (GitHub) · Repository · accessed Oct 8, 2026

  4. 4.
    Center for Advancing Innovation and Standards for Super Intelligence (CAISSI) (external site: nist.gov)

    National Institute of Standards and Technology · Official page · accessed Oct 8, 2026

  5. 5.
    Preparedness Framework, version 2 (external site: cdn.openai.com)

    OpenAI · Official page · published Apr 15, 2025 · accessed Oct 8, 2026

  6. 6.
    Strengthening our Frontier Safety Framework (external site: deepmind.google)

    Google DeepMind · Announcement · published Sep 22, 2025 · accessed Oct 8, 2026

  7. 7.
    Anthropic's Responsible Scaling Policy (external site: anthropic.com)

    Anthropic · Official page · accessed Oct 8, 2026

  8. 8.
    gpt-oss-120b and gpt-oss-20b Model Card (external site: cdn.openai.com)

    OpenAI · Model card · published Aug 5, 2025 · accessed Oct 8, 2026

  9. 9.
    Inspect Petri README (external site: raw.githubusercontent.com)

    Meridian Labs (GitHub) · Repository · accessed Oct 8, 2026

  10. 10.
    Donating our open-source alignment tool (external site: anthropic.com)

    Anthropic · Announcement · published May 7, 2026 · accessed Oct 8, 2026

  11. 11.
    Inspect-Hawk README (external site: raw.githubusercontent.com)

    METR (GitHub) · Repository · accessed Oct 8, 2026

  12. 12.
    AILuminate (external site: mlcommons.org)

    MLCommons · Official page · accessed Oct 8, 2026

  13. 13.
    Llama Guard 4 model card (external site: raw.githubusercontent.com)

    Meta, PurpleLlama (GitHub) · Model card · accessed Oct 8, 2026

  14. 14.
    gpt-oss-safeguard README (external site: raw.githubusercontent.com)

    OpenAI (GitHub) · Repository · accessed Oct 8, 2026

  15. 15.
    gpt-oss-safeguard-20b model card (external site: huggingface.co)

    OpenAI (Hugging Face) · Model card · accessed Oct 8, 2026

  16. 16.
    NIST AI 100-4: Reducing Risks Posed by Synthetic Content (external site: nvlpubs.nist.gov)

    National Institute of Standards and Technology · Official page · published Nov 2024 · accessed Oct 8, 2026

  17. 17.
    C2PA and Content Credentials Explainer (version 2.4) (external site: spec.c2pa.org)

    Coalition for Content Provenance and Authenticity (C2PA) · Documentation · accessed Oct 8, 2026

  18. 18.
    SynthID Text README (external site: raw.githubusercontent.com)

    Google DeepMind (GitHub) · Repository · accessed Oct 8, 2026

  19. 19.
    SynthID Bio README (external site: raw.githubusercontent.com)

    Google DeepMind (GitHub) · Repository · accessed Oct 8, 2026

  20. 20.
    NIST AI 100-2 E2025: Adversarial Machine Learning, A Taxonomy and Terminology of Attacks and Mitigations (external site: nvlpubs.nist.gov)

    National Institute of Standards and Technology · Official page · published Mar 2025 · accessed Oct 8, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project