Independent project. Not a U.S. government website.

USASI
Evaluation toolEvaluation tool

SWE-sweep

Project record

Maintained by Meta123

SWE-sweep is a software-engineering benchmark in which a coding agent is given a real repository containing many known bugs, with no hint about their type or location, and must find and fix as many as it can. Each task is built from real GitHub issue and pull-request pairs at a commit where many of those bugs are present at once, and repairs are graded with hidden tests from the fixing pull requests plus the existing test suite to catch regressions. The paper describes 100 repositories across 22 programming languages.431Fact reviewed Oct 4, 2026

Last reviewedEntry updated Documented release Oct 1, 2026

Availability and license

Overall availability

Public

The harness and task definitions are on GitHub. Running an evaluation requires the Harbor framework and Docker; evaluating a model also requires access to that model and an agent to drive it.14

Availability fact review: Oct 4, 2026

Availability is separate from permission: read the license before using or redistributing.

The MIT license covers the SWE-sweep repository. Tasks are built from third-party open-source repositories that the paper says were selected from permissively licensed projects; each of those projects remains under its own license, which the paper's appendix lists per repository.24Fact reviewed Oct 4, 2026

Component reuse rights

A readable or downloadable component is not automatically reusable. These indicators concern recorded license evidence, not system certification.
weights
Unknown — no complete fact-level rights review
code
Reviewed qualifying license recorded — check scope and conditions
data
Unknown — no complete fact-level rights review
documentation
Unknown — no complete fact-level rights review

No complete system-rights review is recorded for this release.

Public materials checklist

Items for a evaluation tool under USASI rubric v0.2. Unknown means unassessed or insufficient evidence.
Public materials checklist for SWE-sweep
ItemStatusNotes and evidence
CodeIs the evaluation code published?PublicThe evaluation harness is public under the MIT License. The README describes it as a thin wrapper around the Harbor framework.12
Tasks / dataAre the tasks or test data available?PublicTask definitions are in the repository's tasks directory. The paper says the dataset and harness are released, along with rollouts for the models reported in its main results.14
MethodologyIs the method for scoring described?PublicThe paper describes repository selection, task construction, and grading. The README gives the scoring rule: a task scores zero if the patch introduces new failures in the visible test suite; otherwise it scores the share of bugs whose hidden tests pass, and the final score is the share of all bugs resolved across tasks.41
ReproducibilityAre instructions for reproducing results published?PublicThe README gives installation steps, a setup check command, commands to evaluate a patch for one task, and commands to summarize graded runs.1
LimitationsAre known limitations documented?PublicThe paper's discussion states that tasks include only bugs that were reported in an issue and fixed by a pull request, so the benchmark under-counts the bugs actually present in each codebase.4

What it is useful for

Evaluating coding agents on open-ended repository maintenance, where the agent must decide what is broken before repairing it, rather than resolving a single reported issue.43Fact reviewed Oct 4, 2026

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • At release the README warns that Harbor 0.23 does not yet support the separate verifier environments and collect hooks the tasks use, so the project pins a compatible Harbor revision from source until Harbor 0.24 is released. The test suite runs on Python 3.12 and 3.13 in CI.1
  • The paper states that its experiments ran in Docker containers without internet access.4

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The benchmark and harness are published in Meta's facebookresearch GitHub organization, and the MIT license assigns copyright to Meta Platforms, Inc. and affiliates. The project site lists the authors' primary affiliation as Meta Superintelligence Labs, and the paper lists Meta FAIR and another Meta team, with co-authors also affiliated with Harvard University, the University of Washington, and Stanford University. This catalog treats Meta, which holds the repository, as the maintaining entity; see the meta organization record for its U.S. headquarters.1234

Assessed Oct 4, 2026

Sources

  1. 1.
    facebookresearch/swe-sweep repository and README (external site: github.com)

    Meta (GitHub) · Repository · accessed Oct 4, 2026

  2. 2.
  3. 3.
    SWE-sweep (external site: swesweep.com)

    SWE-sweep project · Official page · accessed Oct 4, 2026

  4. 4.
    SWE-sweep: Can Agents Autonomously Find and Fix Bugs? (external site: swesweep.com)

    SWE-sweep project · Paper · published Oct 1, 2026 · accessed Oct 4, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project