WorkflowEvals (TypeSafe AI)
Project record
Maintained by TypeSafe AI1116
WorkflowEvals is TypeSafe AI's evaluation harness for four automation workflows: invoice processing, customer service, agent-trace review, and security-incident triage. Each workflow splits a written policy into narrow yes/no, choice, and score questions for a model, then combines the answers in code into actions. Runs can compare TypeSafe's Jev with models from other providers.25
- Repository: GitHub repository (typesafe-ai/WorkflowEvals) (external site: github.com)
- Website: Workflow evals site (external site: evals.typesafe.ai)
- Dataset hub: WorkflowEvals datasets on Hugging Face (external site: huggingface.co)
- License: License (Apache 2.0) (external site: github.com)
Availability and license
Overall availability
Code is on GitHub under Apache 2.0, and the four workflow datasets download from Hugging Face without login. Running models requires an API key for each provider evaluated, including a TypeSafe API key for the default Jev model.1236
Availability is separate from permission: read the license before using or redistributing.
The README says the datasets are licensed separately in their Hugging Face repositories. On 2026-10-01 the invoice-processing dataset carried an Apache 2.0 license, while the customer-service, security-incidents, and agent-trace-observability dataset repositories stated no license. The Apache 2.0 file in the code repository does not name a copyright holder.2378910
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | Harness code for data loading, model clients, execution, scoring, and plotting is on GitHub under Apache 2.0.123 |
| Tasks / dataAre the tasks or test data available? | Public | Four Hugging Face datasets (150 invoice-processing, 204 customer-service, 111 agent-trace, and 240 security-incident cases) hold inputs, reference labels, and published run results. Only the invoice-processing dataset states a license.2678910 |
| MethodologyIs the method for scoring described? | Public | The evals site and dataset cards describe splitting each policy into Noul, Choice, and Score questions. Consensus reference labels average two providers' large models, and scoring uses exact_actions and primary_action agreement.527 |
| ReproducibilityAre instructions for reproducing results published? | Public | The README documents uv setup, run commands, resuming, pinning a dataset revision, and plotting. Dependency versions are pinned in pyproject.toml. Runs call hosted model APIs and need provider API keys.24 |
| LimitationsAre known limitations documented? | Partial | Dataset cards state that reference labels are model-generated. The evals site says the harness is assumed correct and models are measured against large reference models, so scores measure agreement with those references. The default model is the maintainer's own Jev. No broader discussion of validity was found in the pages read.752 |
What it is useful for
Run and use notes
- The README requires Python 3.13 or later and uv. Supported providers are openai, anthropic, fireworks, groq, cerebras, and typesafe, and each needs its own API key environment variable.2
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
WorkflowEvals is published in the typesafe-ai GitHub organization, which TypeSafe's documentation and launch post link to, and TypeSafe AI's Hugging Face collection links the repository and the evals site. TypeSafe AI, Inc.'s terms of use give its address in San Francisco, California (see the typesafe-ai organization record).11112613
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.