HumanEval
Project record
HumanEval is a code-generation benchmark released by OpenAI with the Codex paper in July 2021. It contains 164 hand-written Python programming problems, each with a function signature, docstring, reference solution, and unit tests, and measures functional correctness of generated code with the pass@k metric. The GitHub repository provides the problem file and the evaluation harness.762
- Repository: GitHub repository (evaluation harness and data) (external site: github.com)
- Dataset hub: Dataset on Hugging Face (openai/openai_humaneval) (external site: huggingface.co)
- Paper: Paper: Evaluating Large Language Models Trained on Code (arXiv 2107.03374) (external site: arxiv.org)
- License: License (MIT) (external site: raw.githubusercontent.com)
Availability and license
Overall availability
Problems and harness are on GitHub under the MIT License, and the problems are also on the Hugging Face Hub without gating. The execution call in the harness is commented out by default; the README asks users to enable it only inside a security sandbox because it runs untrusted model-generated code.126
Availability is separate from permission: read the license before using or redistributing.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | The evaluation harness (Python package human-eval) is in the GitHub repository under the MIT License. The repository changes rarely: its most recent commits (January 2025) fixed a broken evaluation, and the earlier ones date from 2021.1345 |
| Tasks / dataAre the tasks or test data available? | Public | The 164 problems are in data/HumanEval.jsonl.gz in the repository and in the openai/openai_humaneval dataset on Hugging Face (MIT).16 |
| MethodologyIs the method for scoring described? | Public | The paper defines the functional-correctness evaluation and the unbiased pass@k estimator; the README explains that pass@k is not computed when there are fewer samples than k.72 |
| ReproducibilityAre instructions for reproducing results published? | Public | The README documents installation, the expected JSONL sample format, and the evaluation command, with example files for a sanity check.2 |
| LimitationsAre known limitations documented? | Partial | The README warns about executing untrusted code and notes that low memory can cause correct programs to fail. The dataset card says the problems were hand-written to avoid training-set overlap but are likely to appear in later data dumps because they were published on GitHub. A broader discussion of the benchmark's limits was not found in the pages read.26 |
What it is useful for
Checking whether a model's Python completions pass the problem's unit tests, and estimating pass@k from multiple samples per problem with the provided evaluate_functional_correctness command.2
Run and use notes
- The README asks for Python 3.7 or later and a pip install of the cloned repository; the execution call in human_eval/execution.py must be uncommented before evaluation runs.2
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
HumanEval is published in OpenAI's GitHub organization and on OpenAI's Hugging Face organization, and its MIT license names OpenAI as copyright holder. OpenAI Group PBC lists its address as 1455 3rd Street, San Francisco, California, in a February 2026 agreement filed with the SEC.1368
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.