USASI
Evaluation toolEvaluation tool

τ-bench (tau2-bench)

Version 1.0.1

Maintained by Sierra138

τ-bench is Sierra's open-source simulation framework for evaluating customer-service AI agents. In each domain an agent must follow a written policy and use tools while a simulated user takes part in the conversation, in turn-based text mode or full-duplex voice mode. The current repository, which carries the τ³-bench update, covers airline, retail, telecom, and banking-knowledge domains plus a mock domain.28

Last reviewedEntry updated Documented release Jul 22, 2026

Availability and license

Overall availability

Public

Source code and domain task files on GitHub under the MIT License. Running an evaluation requires API access to the language models used for the agent and the simulated user.132

Availability is separate from permission: read the license before using or redistributing.

The domains README (src/tau2/domains/README.md) says each domain's data (policy, task definitions, task splits, and database files) is stored in the same repository under data/tau2/domains/<domain_name>; the LICENSE file does not address data separately.43

Public materials checklist

Items for a evaluation tool under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for τ-bench (tau2-bench)
ItemStatusNotes and evidence
CodeIs the evaluation code published?PublicPublished on GitHub under the MIT License.13
Tasks / dataAre the tasks or test data available?PublicThe domains README documents per-domain data files stored in the repository under data/tau2/domains: policy, task definitions, voice tasks, task splits, and databases.4
MethodologyIs the method for scoring described?PublicThe evaluation guide explains how rewards combine database-state, communication, and assertion checks, and the README cites the τ-bench, τ²-bench, τ-Voice, and τ-Knowledge papers.527
ReproducibilityAre instructions for reproducing results published?PublicThe README documents installation with uv, API-key setup through LiteLLM, and example run commands; versioned releases and a changelog are published, and a pre-v1.0.1 tag is kept for reproducing earlier grading.26
LimitationsAre known limitations documented?PartialThe README states that, after grading fixes to the banking_knowledge domain, results from versions before 1.0.1 are not comparable with later results (other domains are unaffected), and the evaluation guide notes that the reference action sequence is only one valid solution path. No general discussion of the benchmark's validity limits was found in the pages read.25

What it is useful for

Evaluating how well tool-using conversational agents complete customer-service tasks under a domain policy, including tasks in which the user must also take actions, and testing retrieval-based and voice agents.287

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The README requires Python 3.12 or 3.13 (>=3.12, <3.14) and installs with uv; voice features need extra system packages such as portaudio and ffmpeg.2

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The project is maintained in the sierra-research GitHub organization, its MIT license names Sierra Research as copyright holder, and Sierra's own blog says Sierra introduced τ-bench, presents τ²-bench as building on it, and links to this repository. Sierra Technologies, Inc. is a Delaware corporation headquartered in San Francisco, California, per its modern slavery statement.1389

Assessed Sep 29, 2026

Sources

  1. 1.
  2. 2.
    tau2-bench README.md (external site: raw.githubusercontent.com)

    Sierra Research (GitHub) · Documentation · accessed Sep 29, 2026

  3. 3.
    tau2-bench LICENSE (external site: raw.githubusercontent.com)

    Sierra Research (GitHub) · License · accessed Sep 29, 2026

  4. 4.
  5. 5.
    tau2-bench/docs/evaluation.md at main · sierra-research/tau2-bench (external site: github.com)

    Sierra Research (GitHub) · Documentation · accessed Sep 29, 2026

  6. 6.
    Release τ-bench 1.0.1 — banking_knowledge Grading Fixes · sierra-research/tau2-bench (external site: github.com)

    Sierra Research (GitHub) · Release notes · published Jul 22, 2026 · accessed Sep 29, 2026

  7. 7.
  8. 8.
    𝜏²-bench | Sierra (external site: sierra.ai)

    Sierra Technologies, Inc. · Announcement · published Jun 10, 2025 · accessed Sep 29, 2026

  9. 9.
    Modern Slavery Statement | Sierra (external site: sierra.ai)

    Sierra Technologies, Inc. · Official page · published Feb 4, 2026 · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project