Inspect-Hawk (Hawk)
Project record
Maintained by METR (Model Evaluation & Threat Research)12
Inspect-Hawk is METR's open-source platform for running Inspect AI evaluations on cloud infrastructure. Users define tasks, agents, and models in a YAML file; Hawk runs each evaluation in an isolated Kubernetes pod, manages model API credentials through a proxy, streams logs, stores results in a PostgreSQL warehouse, and serves a web interface for browsing them. It is built on Inspect AI, the evaluation framework created by the UK AI Security Institute.13
- Repository: Repository (external site: github.com)
- Documentation: Documentation (external site: hawk.metr.org)
- License: LICENSE (MIT) (external site: github.com)
- Release notes: Releases (external site: github.com)
Availability and license
Overall availability
Source code is public on GitHub and the CLI is installed from PyPI. Running evaluations at scale requires deploying Hawk to the user's own AWS account (EKS, ECS Fargate, Aurora PostgreSQL, S3, and related services) and supplying model provider access.13
Availability is separate from permission: read the license before using or redistributing.
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| Source codeIs the source code publicly readable? | Public | Public repository in the METR GitHub organization.1 |
| DocumentationIs user documentation published? | Public | Documentation site at hawk.metr.org with deployment, configuration, CLI, and evaluation guides.31 |
| InstallationAre installation instructions or packages publicly available? | Public | The README installs the CLI from PyPI with uv or pip (package hawk with the cli extra) and describes deploying an instance with Pulumi, Docker, Python, Node.js, and uv, plus a domain name.1 |
| Supported platformsAre supported operating systems or hardware documented? | Public | Deployment targets AWS; a local mode runs the same evaluation configuration on the user's machine for debugging. The built-in model proxy covers OpenAI, Anthropic, and Google Vertex.13 |
| Release statusAre versioned releases published? | Public | Versioned GitHub releases are published; the latest at review was v3.6.0 (September 24, 2026).4 |
What it is useful for
Teams that run model evaluations regularly and want to run grids of tasks, agents, and models on their own AWS deployment, control which users can run or view results for each model, and run Inspect Scout scans over transcripts from earlier runs.1
Run and use notes
- The README advises keeping the CLI version in sync with the deployment and documents a version check that warns before each command by default.1
Organization context
Provenance and derivatives
Hawk is an original METR repository, not a fork. It depends on Inspect AI, the open-source evaluation framework created by the UK AI Security Institute, and can run scans with Inspect Scout, which is maintained in the meridianlabs-ai GitHub organization.1
- Derived from: Inspect AI (UK AI Security Institute) (external site: inspect.aisi.org.uk) — Evaluation framework Hawk runs; not a catalog member.
U.S. eligibility
Eligible · basis: U.S.-governed project
Hawk is published in METR's GitHub organization, and its MIT license names METR as the 2026 copyright holder. METR (Model Evaluation and Threat Research, Inc.) is listed in the IRS exempt-organization extract for California as a Section 501(c)(3) organization, and its staff work in Berkeley, California (see the METR record). Hawk builds on Inspect AI, which the README attributes to the UK AI Security Institute; that dependency is recorded as provenance and is not a catalog member.125
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.