llm-d
Project record
Maintained by Cloud Native Computing Foundation (Linux Foundation), Sandbox project8910, Red Hat (launched the project; MAINTAINERS.md lists Red Hat staff in project leadership and as community managers)123, Google (MAINTAINERS.md lists Google staff in project leadership)3, IBM (IBM Research is a founding contributor; MAINTAINERS.md lists IBM staff in project leadership and as a community manager)123
llm-d is an open-source distributed inference serving stack for running large language models on Kubernetes. It sits above model servers such as vLLM and adds request routing that is aware of prefix caches and load, KV-cache offloading, prefill/decode disaggregation, expert parallelism for large models, autoscaling, and batch processing. Red Hat launched it in May 2025, and it became a Cloud Native Computing Foundation Sandbox project in March 2026.18129
- Website: Website and documentation (external site: llm-d.ai)
- Repository: Repository (external site: github.com)
- Documentation: Quickstart (external site: llm-d.ai)
- License: LICENSE (Apache 2.0) (external site: github.com)
- Release notes: Releases (external site: github.com)
Availability and license
Overall availability
Source code, container images, and Helm-based guides are public. Deploying a model requires a Kubernetes cluster with accelerators, and the quickstart uses a Hugging Face token stored as a Kubernetes secret.125
Availability is separate from permission: read the license before using or redistributing.
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| Source codeIs the source code publicly readable? | Public | Code is in the llm-d GitHub organization, with experimental components in a separate llm-d-incubation organization.12 |
| DocumentationIs user documentation published? | Public | Documentation and deployment guides are published at llm-d.ai and in the repository.15 |
| InstallationAre installation instructions or packages publicly available? | Public | The quickstart installs the router with Helm and the model server with Kustomize. It needs kubectl, Helm, a Hugging Face token, and the Gateway API Inference Extension CRDs.5 |
| Supported platformsAre supported operating systems or hardware documented? | Public | The accelerator guide names NVIDIA GPUs as the default and also covers AMD ROCm GPUs, Intel XPU, Google Cloud TPUs on GKE, x86_64 CPUs, and several other vendors' accelerators, each with named maintainers. Current images are built on the CUDA 12.9.1 runtime.6 |
| Release statusAre versioned releases published? | Public | Versioned releases are published on GitHub. The latest at review was v0.10.0, published September 29, 2026; the project has not yet reached version 1.0.7 |
What it is useful for
Run and use notes
- PROJECT.md describes the project as building on vLLM and the Kubernetes Gateway API Inference Extension, with code changes made upstream there rather than in forks.2
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
llm-d is a Cloud Native Computing Foundation Sandbox project, accepted on March 12, 2026. The CNCF says it is part of the nonprofit Linux Foundation, whose privacy policy gives a legal postal address in San Francisco, California. Before joining the CNCF, the project was launched by Red Hat, which has its corporate headquarters in Raleigh, North Carolina. Its MAINTAINERS.md lists project leadership employed by Google, IBM, and Red Hat. Day-to-day decisions follow the process in PROJECT.md: lazy consensus, with project maintainers approving process changes. This multi-company project is assessed through its U.S.-based foundation host, not through contributors' nationalities.81011121332
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.