SWE-sweep
Project record
SWE-sweep is a software-engineering benchmark in which a coding agent is given a real repository containing many known bugs, with no hint about their type or location, and must find and fix as many as it can. Each task is built from real GitHub issue and pull-request pairs at a commit where many of those bugs are present at once, and repairs are graded with hidden tests from the fixing pull requests plus the existing test suite to catch regressions. The paper describes 100 repositories across 22 programming languages.431
- Repository: Repository (external site: github.com)
- Website: Project site (external site: swesweep.com)
- Paper: Paper (PDF) (external site: swesweep.com)
- License: LICENSE (MIT) (external site: github.com)
Availability and license
Overall availability
The harness and task definitions are on GitHub. Running an evaluation requires the Harbor framework and Docker; evaluating a model also requires access to that model and an agent to drive it.14
Availability is separate from permission: read the license before using or redistributing.
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | The evaluation harness is public under the MIT License. The README describes it as a thin wrapper around the Harbor framework.12 |
| Tasks / dataAre the tasks or test data available? | Public | Task definitions are in the repository's tasks directory. The paper says the dataset and harness are released, along with rollouts for the models reported in its main results.14 |
| MethodologyIs the method for scoring described? | Public | The paper describes repository selection, task construction, and grading. The README gives the scoring rule: a task scores zero if the patch introduces new failures in the visible test suite; otherwise it scores the share of bugs whose hidden tests pass, and the final score is the share of all bugs resolved across tasks.41 |
| ReproducibilityAre instructions for reproducing results published? | Public | The README gives installation steps, a setup check command, commands to evaluate a patch for one task, and commands to summarize graded runs.1 |
| LimitationsAre known limitations documented? | Public | The paper's discussion states that tasks include only bugs that were reported in an issue and fixed by a pull request, so the benchmark under-counts the bugs actually present in each codebase.4 |
What it is useful for
Run and use notes
- At release the README warns that Harbor 0.23 does not yet support the separate verifier environments and collect hooks the tasks use, so the project pins a compatible Harbor revision from source until Harbor 0.24 is released. The test suite runs on Python 3.12 and 3.13 in CI.1
- The paper states that its experiments ran in Docker containers without internet access.4
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
The benchmark and harness are published in Meta's facebookresearch GitHub organization, and the MIT license assigns copyright to Meta Platforms, Inc. and affiliates. The project site lists the authors' primary affiliation as Meta Superintelligence Labs, and the paper lists Meta FAIR and another Meta team, with co-authors also affiliated with Harvard University, the University of Washington, and Stanford University. This catalog treats Meta, which holds the repository, as the maintaining entity; see the meta organization record for its U.S. headquarters.1234
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.