USASI
Evaluation toolEvaluation tool

HumanEval

Project record

Maintained by OpenAI136

HumanEval is a code-generation benchmark released by OpenAI with the Codex paper in July 2021. It contains 164 hand-written Python programming problems, each with a function signature, docstring, reference solution, and unit tests, and measures functional correctness of generated code with the pass@k metric. The GitHub repository provides the problem file and the evaluation harness.762

Last reviewedEntry updated Documented release Jul 2021

Availability and license

Overall availability

Public

Problems and harness are on GitHub under the MIT License, and the problems are also on the Hugging Face Hub without gating. The execution call in the harness is commented out by default; the README asks users to enable it only inside a security sandbox because it runs untrusted model-generated code.126

Availability is separate from permission: read the license before using or redistributing.

Public materials checklist

Items for a evaluation tool under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for HumanEval
ItemStatusNotes and evidence
CodeIs the evaluation code published?PublicThe evaluation harness (Python package human-eval) is in the GitHub repository under the MIT License. The repository changes rarely: its most recent commits (January 2025) fixed a broken evaluation, and the earlier ones date from 2021.1345
Tasks / dataAre the tasks or test data available?PublicThe 164 problems are in data/HumanEval.jsonl.gz in the repository and in the openai/openai_humaneval dataset on Hugging Face (MIT).16
MethodologyIs the method for scoring described?PublicThe paper defines the functional-correctness evaluation and the unbiased pass@k estimator; the README explains that pass@k is not computed when there are fewer samples than k.72
ReproducibilityAre instructions for reproducing results published?PublicThe README documents installation, the expected JSONL sample format, and the evaluation command, with example files for a sanity check.2
LimitationsAre known limitations documented?PartialThe README warns about executing untrusted code and notes that low memory can cause correct programs to fail. The dataset card says the problems were hand-written to avoid training-set overlap but are likely to appear in later data dumps because they were published on GitHub. A broader discussion of the benchmark's limits was not found in the pages read.26

What it is useful for

Checking whether a model's Python completions pass the problem's unit tests, and estimating pass@k from multiple samples per problem with the provided evaluate_functional_correctness command.2

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The README asks for Python 3.7 or later and a pip install of the cloned repository; the execution call in human_eval/execution.py must be uncommented before evaluation runs.2

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

HumanEval is published in OpenAI's GitHub organization and on OpenAI's Hugging Face organization, and its MIT license names OpenAI as copyright holder. OpenAI Group PBC lists its address as 1455 3rd Street, San Francisco, California, in a February 2026 agreement filed with the SEC.1368

Assessed Sep 29, 2026

Sources

  1. 1.
    openai/human-eval (external site: github.com)

    OpenAI · Repository · accessed Sep 29, 2026

  2. 2.
    human-eval README.md (external site: raw.githubusercontent.com)

    OpenAI · Documentation · accessed Sep 29, 2026

  3. 3.
    human-eval LICENSE (external site: raw.githubusercontent.com)

    OpenAI · License · accessed Sep 29, 2026

  4. 4.
    human-eval setup.py (external site: raw.githubusercontent.com)

    OpenAI · Repository · accessed Sep 29, 2026

  5. 5.
    Commits · openai/human-eval (external site: github.com)

    OpenAI · Repository · accessed Sep 29, 2026

  6. 6.
    openai/openai_humaneval dataset card (external site: huggingface.co)

    OpenAI · Dataset card · accessed Sep 29, 2026

  7. 7.
    Evaluating Large Language Models Trained on Code (external site: arxiv.org)

    arXiv (OpenAI authors) · Paper · published Jul 7, 2021 · accessed Sep 29, 2026

  8. 8.
    Exhibit 10.1: Equity commitment letter agreement between OpenAI Group PBC and Amazon (external site: sec.gov)

    U.S. Securities and Exchange Commission (Amazon.com, Inc. filing) · Filing · published Feb 27, 2026 · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project