Independent project. Not a U.S. government website.

USASI
Evaluation toolEvaluation tool

WorkflowEvals (TypeSafe AI)

Project record

Maintained by TypeSafe AI1116

WorkflowEvals is TypeSafe AI's evaluation harness for four automation workflows: invoice processing, customer service, agent-trace review, and security-incident triage. Each workflow splits a written policy into narrow yes/no, choice, and score questions for a model, then combines the answers in code into actions. Runs can compare TypeSafe's Jev with models from other providers.25Fact reviewed Oct 1, 2026

Last reviewedEntry updated Documented release Unknown

Availability and license

Overall availability

Public

Code is on GitHub under Apache 2.0, and the four workflow datasets download from Hugging Face without login. Running models requires an API key for each provider evaluated, including a TypeSafe API key for the default Jev model.1236

Availability fact review: Oct 1, 2026

Availability is separate from permission: read the license before using or redistributing.

The README says the datasets are licensed separately in their Hugging Face repositories. On 2026-10-01 the invoice-processing dataset carried an Apache 2.0 license, while the customer-service, security-incidents, and agent-trace-observability dataset repositories stated no license. The Apache 2.0 file in the code repository does not name a copyright holder.2378910Fact reviewed Oct 1, 2026

Component reuse rights

A readable or downloadable component is not automatically reusable. These indicators concern recorded license evidence, not system certification.
weights
Unknown — no complete fact-level rights review
code
Reviewed qualifying license recorded — check scope and conditions
data
Unknown — no complete fact-level rights review
documentation
Unknown — no complete fact-level rights review

No complete system-rights review is recorded for this release.

Public materials checklist

Items for a evaluation tool under USASI rubric v0.2. Unknown means unassessed or insufficient evidence.
Public materials checklist for WorkflowEvals (TypeSafe AI)
ItemStatusNotes and evidence
CodeIs the evaluation code published?PublicHarness code for data loading, model clients, execution, scoring, and plotting is on GitHub under Apache 2.0.123
Tasks / dataAre the tasks or test data available?PublicFour Hugging Face datasets (150 invoice-processing, 204 customer-service, 111 agent-trace, and 240 security-incident cases) hold inputs, reference labels, and published run results. Only the invoice-processing dataset states a license.2678910
MethodologyIs the method for scoring described?PublicThe evals site and dataset cards describe splitting each policy into Noul, Choice, and Score questions. Consensus reference labels average two providers' large models, and scoring uses exact_actions and primary_action agreement.527
ReproducibilityAre instructions for reproducing results published?PublicThe README documents uv setup, run commands, resuming, pinning a dataset revision, and plotting. Dependency versions are pinned in pyproject.toml. Runs call hosted model APIs and need provider API keys.24
LimitationsAre known limitations documented?PartialDataset cards state that reference labels are model-generated. The evals site says the harness is assumed correct and models are measured against large reference models, so scores measure agreement with those references. The default model is the maintainer's own Jev. No broader discussion of validity was found in the pages read.752

What it is useful for

Comparing models on structured decision tasks inside fixed workflows, by agreement with reference decisions, cost per case, and time per case, and comparing a workflow approach with sending the same policy as a single prompt.25Fact reviewed Oct 1, 2026

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The README requires Python 3.13 or later and uv. Supported providers are openai, anthropic, fireworks, groq, cerebras, and typesafe, and each needs its own API key environment variable.2

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

WorkflowEvals is published in the typesafe-ai GitHub organization, which TypeSafe's documentation and launch post link to, and TypeSafe AI's Hugging Face collection links the repository and the evals site. TypeSafe AI, Inc.'s terms of use give its address in San Francisco, California (see the typesafe-ai organization record).11112613

Assessed Oct 1, 2026

Sources

  1. 1.
    typesafe-ai/WorkflowEvals (GitHub repository) (external site: github.com)

    TypeSafe AI, Inc. (GitHub) · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  2. 2.
    WorkflowEvals README.md (external site: raw.githubusercontent.com)

    TypeSafe AI, Inc. (GitHub) · Documentation · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  3. 3.
    WorkflowEvals LICENSE (Apache License 2.0) (external site: raw.githubusercontent.com)

    TypeSafe AI, Inc. (GitHub) · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  4. 4.
    WorkflowEvals pyproject.toml (external site: raw.githubusercontent.com)

    TypeSafe AI, Inc. (GitHub) · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  5. 5.
    Workflow evals - TypeSafe AI (external site: evals.typesafe.ai)

    TypeSafe AI, Inc. · Official page · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  6. 6.
    WorkflowEvals collection (TypeSafe AI on Hugging Face) (external site: huggingface.co)

    TypeSafe AI, Inc. (Hugging Face) · Dataset card · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  7. 7.
    typesafe/evalsafe-invoice-processing dataset card (external site: huggingface.co)

    TypeSafe AI, Inc. (Hugging Face) · Dataset card · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  8. 8.
    typesafe/evalsafe-customer-service dataset card (external site: huggingface.co)

    TypeSafe AI, Inc. (Hugging Face) · Dataset card · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  9. 9.
    typesafe/evalsafe-security-incidents dataset card (external site: huggingface.co)

    TypeSafe AI, Inc. (Hugging Face) · Dataset card · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  10. 10.
    typesafe/evalsafe-agent-trace-observability dataset card (external site: huggingface.co)

    TypeSafe AI, Inc. (Hugging Face) · Dataset card · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  11. 11.
    TypeSafe (typesafe-ai) on GitHub (external site: github.com)

    TypeSafe AI, Inc. (GitHub) · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  12. 12.
    Introducing System One Models & Jev - TypeSafe AI Blog (external site: typesafe.ai)

    TypeSafe AI, Inc. · Announcement · published Sep 15, 2026 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  13. 13.
    Terms of use - TypeSafe AI (external site: typesafe.ai)

    TypeSafe AI, Inc. · Official page · published Sep 19, 2026 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project