USASI
Evaluation toolEvaluation tool

LM Evaluation Harness

Version 0.4.13

Maintained by EleutherAI125

An open-source framework from EleutherAI for evaluating language models on many benchmark tasks through one interface. It supports local models (for example through Hugging Face Transformers or vLLM) and hosted model APIs, and its README lists more than 60 standard academic benchmarks with hundreds of subtasks and variants.1

Last reviewedEntry updated Documented release Aug 31, 2026

Availability and license

Overall availability

Public

Source code on GitHub under the MIT License; installable from PyPI as lm-eval, with model backends installed as optional extras.152

Availability is separate from permission: read the license before using or redistributing.

The new-task guide states that task data is downloaded and managed through the Hugging Face datasets API, so benchmark data comes from separately published datasets rather than from the harness repository.3

Public materials checklist

Items for a evaluation tool under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for LM Evaluation Harness
ItemStatusNotes and evidence
CodeIs the evaluation code published?PublicPublished on GitHub under the MIT License.12
Tasks / dataAre the tasks or test data available?PublicTask configurations are in the repository; task data is loaded from datasets on the Hugging Face Hub (or local files) as specified in each task's configuration.31
MethodologyIs the method for scoring described?PublicThe new-task guide describes generative and multiple-choice (log-likelihood) task types and how metrics and aggregations are declared; the README sets out the order of preference used to choose prompting and evaluation procedures.31
ReproducibilityAre instructions for reproducing results published?PublicThe README documents command-line usage, publicly available prompts, result caching, and logging; versioned releases are published on GitHub and PyPI.145
LimitationsAre known limitations documented?PartialThe README documents operational limitations (no native multi-node evaluation, early-stage support for the PyTorch MPS backend, and request types some backends do not support). A general discussion of benchmark validity limits was not found in the pages read.1

What it is useful for

Running benchmark evaluations of language models with shared, publicly available prompts, and adding new tasks through YAML task configurations.13

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • Since December 2025 the base package no longer bundles transformers or torch; backends are installed as extras such as lm_eval[hf], lm_eval[vllm], or lm_eval[api], per the README.1

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The project is maintained in EleutherAI's GitHub organization, its MIT license names EleutherAI as copyright holder, and the PyPI package lists EleutherAI as author. EleutherAI Institute is listed by the IRS as a 501(c)(3) organization in Washington, D.C.1256

Assessed Sep 29, 2026

Sources

  1. 1.
    EleutherAI/lm-evaluation-harness (external site: github.com)

    EleutherAI · Repository · accessed Sep 29, 2026

  2. 2.
  3. 3.
  4. 4.
    Release v0.4.13 · EleutherAI/lm-evaluation-harness (external site: github.com)

    EleutherAI · Release notes · published Aug 31, 2026 · accessed Sep 29, 2026

  5. 5.
    lm-eval on PyPI (external site: pypi.org)

    Python Package Index · Documentation · published Aug 31, 2026 · accessed Sep 29, 2026

  6. 6.

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project