Hub
Benchmarks and evaluation
Evaluation harnesses, benchmarks, and the organizations that run them, and what a published result can and cannot tell you about who tested what, on which version, and under which conditions.
Evaluation records in this catalog are of three kinds. Harnesses run many tests through one interface: EleutherAI's lm-evaluation-harness tests generative language models on a large number of tasks, using local models or commercial APIs, and Stanford CRFM's HELM offers benchmarks in a standardized format, models from several providers behind one interface, and metrics beyond accuracy such as efficiency, bias, and toxicity. Benchmarks are the task sets themselves, such as Humanity's Last Exam, a set of multiple-choice and short-answer questions suited to automated grading. System benchmarks such as MLCommons' MLPerf measure how fast computer systems train a model to a target quality or process inputs with a trained model.1246
Who produced a result matters. Under MLPerf's submission rules, each submitter must review at least one other submission, and the submitted source code must be sufficient to reproduce the results. ARC Prize Foundation tests models for its Verified Leaderboard on a semi-private evaluation set, working with model providers it selects, while Community Leaderboard entries get a light review and are not verified by default.78
Benchmarks change over time, so a result belongs to a version. Terminal-Bench uses semantic versioning: a change that significantly alters the agent's environment is a major version that requires re-running agents, while a change to the verifier alone is minor and lets saved runs be re-graded, and each leaderboard corresponds to a specific dataset version. Humanity's Last Exam keeps a dated public log of questions added, removed, updated, and re-added. HELM entered maintenance mode on June 1, 2026: no new evaluations will be added to its leaderboards, and its maintainers warn that scenarios and models that depend on external APIs may break as those APIs change or providers deprecate models.953
How a model is reached limits what can be measured. In lm-evaluation-harness, models that do not return log probabilities, such as those behind chat-completion APIs, can run only generation tasks, while local models can run every task type; its README suggests checking a few sample outputs to confirm that answers are extracted and scored as intended. Test material can also reach training data. Humanity's Last Exam includes a canary string so model builders can filter it out, and ARC Prize Foundation acknowledges possible limited leakage over time because its tasks are sent to external APIs.148
This catalog does not publish the scores that benchmarks produce. Each evaluation record instead checks whether the evaluation code and the tasks or test data are available, whether the scoring method is described, whether instructions for reproducing results are published, and whether known limitations are documented.10
What this hub covers
Covers evaluation harnesses, benchmarks, safety and system-performance tests, and the organizations that maintain them. It explains how results are produced and where they stop; it does not report scores, rank models, or say which benchmark to trust. Featured records are examples chosen to cover the subject, not a ranking or a complete list.
Reading path
- Reading evaluationsStart here for the questions to ask of any published result.
- BenchmarkA fixed set of tasks with a scoring method, and what a reported result is.
- Evaluation harnessSoftware that runs benchmarks the same way each time so others can re-run them.
- Source standardsHow the catalog treats reported results, and why it does not publish scores.
- How to read a model cardModel cards report evaluations; check what was tested and how before relying on it.
- AI agents and agent toolingAgent benchmarks test a model together with the software wrapped around it.
- Chips, cloud, and computeWhere system benchmarks such as MLPerf fit among hardware and cloud providers.
In the catalog
- Evaluation tools and benchmarksEach record has a checklist covering code, tasks, scoring, and reproducibility.
- Records that mention benchmarksA text search that also finds datasets and tools used in evaluation.
- Standards bodiesOrganizations that publish benchmarks, measurement methods, or standards.
Featured records
Examples chosen to cover the subject; not a ranking or a complete list.
- MLCommonsOrganization
- ARC Prize FoundationOrganization
- Center for AI Safety (CAIS)Organization
- LM Evaluation HarnessOpen model or tool
- HELM (Holistic Evaluation of Language Models)Open model or tool
- Terminal-BenchOpen model or tool
- MLPerfOpen model or tool
- ARC-AGIOpen model or tool
- Humanity's Last Exam (HLE)Open model or tool
- AILuminateOpen model or tool
- τ-bench (tau2-bench)Open model or tool
Primary documents
- General MLPerf Submission Rules (external site: raw.githubusercontent.com)MLCommons (GitHub)MLPerf's submission rules, including peer review between submitters and the requirement that submitted code can reproduce the results.
- ARC Prize Verified Testing Policy (external site: arcprize.org)ARC Prize FoundationARC Prize Foundation's testing policy, which separates verified results from community submissions and discusses leakage through external APIs.
- Continuous Benchmarks (external site: tbench.ai)Terminal-BenchExplains how a benchmark can be versioned like software, and which changes require re-running or only re-grading.
- HLE rolling changes log (external site: raw.githubusercontent.com)Center for AI Safety (GitHub)A dated log of question changes, which shows that a benchmark's contents can differ between two results.
- Maintenance Mode Policy (HELM documentation) (external site: crfm-helm.readthedocs.io)Stanford CRFMWhat maintenance mode means for a harness whose tests depend on external model APIs.
- Language Model Evaluation Harness README (external site: raw.githubusercontent.com)EleutherAI (GitHub)The harness's own description of supported model types, and of which task types need log probabilities.
- Methodology (external site: unitedstatesofamericasuperintelligence.com)USASIUSASI's source standards, including its rule against publishing benchmark results.
Sources · reviewed Oct 2, 2026
Support Us
Help keep USASI useful.
Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).
About supporting this project