USASI
Evaluation toolEvaluation tool

Humanity's Last Exam (HLE)

Version HLE (2,500 questions, finalized April 2025); HLE-Rolling; HLE-Diamond (September 2026)

Maintained by Center for AI Safety1239, Scale AI19

A multimodal benchmark of 2,500 expert-written, closed-ended academic questions across more than a hundred subjects, including mathematics, the humanities, and the natural sciences, in multiple-choice and short-answer formats suited to automated grading. It was organized by teams at the Center for AI Safety and Scale AI with questions from nearly 1,000 subject-matter contributors, and was published in Nature in January 2026.12

Last reviewedEntry updated Documented release Jan 2025

Availability and license

Overall availability

Partial

The public question set is on Hugging Face behind an automatic access gate that asks users to share contact information; the evaluation code is public on GitHub. A private held-out question set is not released.4521

Availability is separate from permission: read the license before using or redistributing.

The dataset and repository include a canary string intended to help model developers keep the benchmark out of training data. The separate cais/hle-diamond dataset's metadata also lists the MIT license and uses the same automatic access gate.246

Public materials checklist

Items for a evaluation tool under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for Humanity's Last Exam (HLE)
ItemStatusNotes and evidence
CodeIs the evaluation code published?PublicThe hle_eval scripts generate model predictions through the openai-python interface and grade them with a judge model; MIT license.23
Tasks / dataAre the tasks or test data available?PartialThe 2,500 public questions are downloadable after accepting the Hugging Face access prompt; a private held-out set is kept back to assess overfitting.451
MethodologyIs the method for scoring described?PublicThe paper (arXiv, with a Nature version) describes the benchmark, whose questions each have a known, unambiguous, verifiable answer that cannot be quickly found by internet search. The website explains that calibration error is measured by asking models for an answer and a 0-100% confidence, with answers graded by a judge model.81
ReproducibilityAre instructions for reproducing results published?PublicThe README gives commands for running predictions and the judge, notes temperature 0 as the default, and advises at least 8,192 completion tokens for reasoning models.2
LimitationsAre known limitations documented?PublicThe website states that HLE tests structured academic problems rather than open-ended research or creative problem solving, and that high accuracy alone would not indicate autonomous research ability. Changes to the question set (removals, re-additions, updates, and additions) are listed in the public HLE-Rolling change log, and HLE-Diamond is described as a subset refined over a year of cleaning.1109

What it is useful for

Measuring language model accuracy and calibration on difficult closed-ended academic questions; the organizers keep a private held-out question set to check for overfitting.12

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • HLE-Rolling (October 2025) is a dynamic fork updated through a public change log, and HLE-Diamond (September 22, 2026) is a refined 1,000-question subset split evenly between reasoning and knowledge questions, published as cais/hle-diamond.1910

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The benchmark's organizing team is drawn from the Center for AI Safety and Scale AI; its GitHub repository is in the centerforaisafety organization under an MIT license held by centerforaisafety, and its dataset is published under CAIS's Hugging Face organization. The Center for AI Safety is a San Francisco 501(c)(3) nonprofit listed by the IRS. Scale AI is covered by its own catalog record.12371112

Assessed Sep 29, 2026

Sources

  1. 1.
    Humanity's Last Exam (external site: lastexam.ai)

    Center for AI Safety and Scale AI · Official page · accessed Sep 29, 2026

  2. 2.
    centerforaisafety/hle README (external site: github.com)

    Center for AI Safety · Repository · accessed Sep 29, 2026

  3. 3.
    centerforaisafety/hle LICENSE (external site: raw.githubusercontent.com)

    Center for AI Safety · License · accessed Sep 29, 2026

  4. 4.
    cais/hle dataset card (external site: huggingface.co)

    Center for AI Safety · Dataset card · accessed Sep 29, 2026

  5. 5.
  6. 6.
  7. 7.
  8. 8.
    Humanity's Last Exam (arXiv 2501.14249) (external site: arxiv.org)

    arXiv · Paper · published Jan 24, 2025 · accessed Sep 29, 2026

  9. 9.
    HLE-Diamond release notes (external site: lastexam.ai)

    Center for AI Safety and Scale AI · Release notes · published Sep 22, 2026 · accessed Sep 29, 2026

  10. 10.
    hle-rolling-changes.txt (centerforaisafety/hle) (external site: github.com)

    Center for AI Safety · Release notes · accessed Sep 29, 2026

  11. 11.
    Donate | CAIS (external site: safe.ai)

    Center for AI Safety · Official page · accessed Sep 29, 2026

  12. 12.

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project