Harvey LAB (Legal Agent Benchmark)
Project record
LAB is an open-source benchmark from Harvey for evaluating AI agents on legal work. It consists of a task set, in which each task gives an agent instructions and a matter file of documents and defines a rubric of pass/fail criteria, and an execution harness for running agents against the tasks and grading their work product. The tasks span transactional, advisory, regulatory, and litigation work across many legal practice areas.183
- Repository: Repository (external site: github.com)
- Release notes: Announcement (external site: harvey.ai)
- Documentation: Evaluation methodology (external site: github.com)
- License: LICENSE (MIT) (external site: github.com)
Availability and license
Overall availability
Tasks and harness are public on GitHub. Running the benchmark requires API access to the model under test and to the LLM judge models, which are used under their providers' terms.13
Availability is separate from permission: read the license before using or redistributing.
Component reuse rights
- weights
- Reviewed qualifying license recorded — check scope and conditions
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Reviewed qualifying license recorded — check scope and conditions
- documentation
- Reviewed qualifying license recorded — check scope and conditions
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | The harness and grader are in the repository and are also built as the lab-core Python package, attached as a wheel to each GitHub release.1 |
| Tasks / dataAre the tasks or test data available? | Public | Task instructions, matter documents, and rubrics are in the repository's tasks directory, organized by practice area.14 |
| MethodologyIs the method for scoring described? | Public | Each rubric criterion is graded pass or fail by LLM judges (by default one Anthropic and one OpenAI model) reading only the deliverables relevant to that criterion; a task scores 1 only if every criterion passes, and the per-criterion pass rate is reported as a diagnostic.3 |
| ReproducibilityAre instructions for reproducing results published? | Public | A tutorial walks through setup, running an agent on one task, scoring, and comparing runs; results also depend on the judge models chosen.43 |
| LimitationsAre known limitations documented? | Partial | The tutorial says the matter documents were generated synthetically under the guidance and review of lawyers and contain imperfections. The announcement describes LAB as an ongoing project and says it launched without a leaderboard.48 |
What it is useful for
Run and use notes
Organization context
Provenance and derivatives
U.S. eligibility
Eligible · basis: U.S.-governed project
LAB is published in Harvey's harveyai GitHub organization, and its MIT license names Harvey AI as the 2026 copyright holder. Harvey's privacy policy names Harvey AI Corporation in San Francisco, California, as its primary entity (see the Harvey record). No outside governance is documented.129
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.