Independent project. Not a U.S. government website.

USASI

Hub

Benchmarks and evaluation

Evaluation harnesses, benchmarks, and the organizations that run them, and what a published result can and cannot tell you about who tested what, on which version, and under which conditions.

Evaluation records in this catalog are of three kinds. Harnesses run many tests through one interface: EleutherAI's lm-evaluation-harness tests generative language models on a large number of tasks, using local models or commercial APIs, and Stanford CRFM's HELM offers benchmarks in a standardized format, models from several providers behind one interface, and metrics beyond accuracy such as efficiency, bias, and toxicity. Benchmarks are the task sets themselves, such as Humanity's Last Exam, a set of multiple-choice and short-answer questions suited to automated grading. System benchmarks such as MLCommons' MLPerf measure how fast computer systems train a model to a target quality or process inputs with a trained model.1246

Who produced a result matters. Under MLPerf's submission rules, each submitter must review at least one other submission, and the submitted source code must be sufficient to reproduce the results. ARC Prize Foundation tests models for its Verified Leaderboard on a semi-private evaluation set, working with model providers it selects, while Community Leaderboard entries get a light review and are not verified by default.78

Benchmarks change over time, so a result belongs to a version. Terminal-Bench uses semantic versioning: a change that significantly alters the agent's environment is a major version that requires re-running agents, while a change to the verifier alone is minor and lets saved runs be re-graded, and each leaderboard corresponds to a specific dataset version. Humanity's Last Exam keeps a dated public log of questions added, removed, updated, and re-added. HELM entered maintenance mode on June 1, 2026: no new evaluations will be added to its leaderboards, and its maintainers warn that scenarios and models that depend on external APIs may break as those APIs change or providers deprecate models.953

How a model is reached limits what can be measured. In lm-evaluation-harness, models that do not return log probabilities, such as those behind chat-completion APIs, can run only generation tasks, while local models can run every task type; its README suggests checking a few sample outputs to confirm that answers are extracted and scored as intended. Test material can also reach training data. Humanity's Last Exam includes a canary string so model builders can filter it out, and ARC Prize Foundation acknowledges possible limited leakage over time because its tasks are sent to external APIs.148

This catalog does not publish the scores that benchmarks produce. Each evaluation record instead checks whether the evaluation code and the tasks or test data are available, whether the scoring method is described, whether instructions for reproducing results are published, and whether known limitations are documented.10

What this hub covers

Covers evaluation harnesses, benchmarks, safety and system-performance tests, and the organizations that maintain them. It explains how results are produced and where they stop; it does not report scores, rank models, or say which benchmark to trust. Featured records are examples chosen to cover the subject, not a ranking or a complete list.

Reading path

  1. 1.Reading evaluationsStart here for the questions to ask of any published result.
  2. 2.BenchmarkA fixed set of tasks with a scoring method, and what a reported result is.
  3. 3.Evaluation harnessSoftware that runs benchmarks the same way each time so others can re-run them.
  4. 4.Source standardsHow the catalog treats reported results, and why it does not publish scores.
  5. 5.How to read a model cardModel cards report evaluations; check what was tested and how before relying on it.
  6. 6.AI agents and agent toolingAgent benchmarks test a model together with the software wrapped around it.
  7. 7.Chips, cloud, and computeWhere system benchmarks such as MLPerf fit among hardware and cloud providers.

In the catalog

Examples chosen to cover the subject; not a ranking or a complete list.

Primary documents

Sources · reviewed Oct 2, 2026

  1. 1.
    Language Model Evaluation Harness README (external site: raw.githubusercontent.com)

    EleutherAI (GitHub) · Repository · accessed Oct 2, 2026

  2. 2.
  3. 3.
  4. 4.
    Humanity's Last Exam README (external site: raw.githubusercontent.com)

    Center for AI Safety (GitHub) · Repository · accessed Oct 2, 2026

  5. 5.
    HLE rolling changes log (external site: raw.githubusercontent.com)

    Center for AI Safety (GitHub) · Release notes · accessed Oct 2, 2026

  6. 6.
    Benchmarks (external site: mlcommons.org)

    MLCommons · Official page · accessed Oct 2, 2026

  7. 7.
    General MLPerf Submission Rules (external site: raw.githubusercontent.com)

    MLCommons (GitHub) · Documentation · accessed Oct 2, 2026

  8. 8.
    ARC Prize Verified Testing Policy (external site: arcprize.org)

    ARC Prize Foundation · Official page · accessed Oct 2, 2026

  9. 9.
    Continuous Benchmarks (external site: tbench.ai)

    Terminal-Bench · Announcement · accessed Oct 2, 2026

  10. 10.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project