BigCodeBench
Project record
Maintained by BigCode project (open scientific collaboration)418, Hugging Face (co-leads the BigCode steering committee)9, ServiceNow (co-leads the BigCode steering committee)9
BigCodeBench is a Python code-generation benchmark from the BigCode project with 1,140 function-level tasks that require calls to 139 libraries across 7 domains. It has two variants, Complete (prompts with structured docstrings) and Instruct (natural-language instructions for chat models), plus a 148-task Hard subset, and it is scored by executing unit tests. The GitHub repository was archived and made read-only on July 20, 2026.5241
- Repository: GitHub repository (archived July 2026) (external site: github.com)
- Dataset hub: Dataset on Hugging Face (bigcode/bigcodebench) (external site: huggingface.co)
- Paper: Paper: BigCodeBench (arXiv 2406.15877) (external site: arxiv.org)
- License: License (Apache 2.0) (external site: raw.githubusercontent.com)
Availability and license
Overall availability
Tasks are on the Hugging Face Hub without gating, and the evaluation package is on PyPI as bigcodebench (latest version 0.2.5). The GitHub repository was archived on July 20, 2026 and is read-only.461
Availability is separate from permission: read the license before using or redistributing.
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Unknown — no complete fact-level rights review
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | Evaluation code on GitHub (archived, read-only) under Apache 2.0 and on PyPI; the last GitHub release is v0.2.5 (April 2025, marked as a pre-release).1367 |
| Tasks / dataAre the tasks or test data available? | Public | The Hugging Face dataset holds the tasks with prompts, reference solutions, unit tests, and library lists, published in versions v0.1.0 through v0.1.4.4 |
| MethodologyIs the method for scoring described? | Public | The paper describes task construction and evaluation by unit-test execution; the README explains the Complete and Instruct splits and the separate prompting of base and chat models.52 |
| ReproducibilityAre instructions for reproducing results published? | Public | The README documents installation and the bigcodebench.evaluate command, warns that batch inference can change greedy-decoding results, and links pre-generated model samples attached to release v0.2.4.2 |
| LimitationsAre known limitations documented? | Public | The dataset card notes that tasks cover only English and Python, lists limitation areas (multilingualism, saturation, reliability, efficiency, rigorousness, generalization, evolution, and interaction), and refers to Appendix D of the technical report for details.4 |
What it is useful for
Evaluating code completion and instruction-following code generation that involves multiple library calls, with local, E2B-sandbox, or remote execution of the generated code.2
Run and use notes
- The README states that evaluation uses a remote execution API by default (a Hugging Face Space) and also supports local or E2B-sandbox execution, the latter requiring an E2B API key.2
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
Multi-organization assessment. BigCodeBench's dataset card, published in the BigCode Hugging Face organization, says the dataset was created as part of the BigCode Project, and its code and leaderboard are published under BigCode's GitHub and Hugging Face organizations. BigCode's organization page says the project is jointly led by Hugging Face and ServiceNow and is governed by a steering committee led by the two companies, which organizes and manages the project, oversees all working groups, and makes tie-breaking decisions as a last resort; the mission page says technical governance otherwise takes place in community working groups. Both companies are recorded in this catalog as U.S.-headquartered (see their records), so the documented governing entities are U.S.-based; the BigCode StarCoder2 records rest on the same governance. The dataset's lead curator lists Monash University and CSIRO's Data61 in Australia, and other contributors come from many institutions; contributor affiliations do not affect eligibility. The GitHub repository has been archived and read-only since July 2026.4198
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.