Independent project. Not a U.S. government website.

USASI
Model release

NVIDIA parakeet-unified-en-0.6b

Release in the NVIDIA Parakeet family · version parakeet-unified-en-0.6b

Maintained by NVIDIA (NeMo)16

parakeet-unified-en-0.6b is a 600M-parameter English speech recognition model that serves both offline and streaming transcription from one checkpoint. It pairs a 24-layer FastConformer encoder, trained jointly in offline and streaming modes with a mode-consistency regularization loss, with an RNN-T decoder; streaming latency can be set from 2,080 ms down to 160 ms. Output includes punctuation and capitalization. NVIDIA released it on Hugging Face on April 7, 2026.13Fact reviewed Oct 1, 2026

Last reviewedEntry updated Documented release Apr 7, 2026

Availability and license

Overall availability

Public

The .nemo checkpoint is downloadable from Hugging Face without an access gate under the NVIDIA Open Model License Agreement; the card states the model is ready for commercial and non-commercial use and lists global deployment geography.21

Availability fact review: Oct 1, 2026

Availability is separate from permission: read the license before using or redistributing.

The NVIDIA Open Model License Agreement (last modified October 24, 2025) is NVIDIA's own license. It grants a perpetual, worldwide, royalty-free license, revocable on the conditions it states, to use, modify, and distribute the model; it states that models are commercially usable and that NVIDIA claims no ownership of outputs, and lets licensees own their derivative models. Redistributors must include a copy of the agreement and a "Licensed by NVIDIA Corporation under the NVIDIA Open Model License" notice. Rights end for anyone who brings copyright or patent litigation over the model, or who bypasses its safety guardrails without a substitute. Use must follow NVIDIA's Trustworthy AI terms, users indemnify NVIDIA against third-party claims, and NVIDIA may update the agreement. This differs from the CC-BY-4.0 terms of parakeet-tdt-0.6b-v3.41Fact reviewed Oct 1, 2026

Component reuse rights

A readable or downloadable component is not automatically reusable. These indicators concern recorded license evidence, not system certification.
weights
Unknown — no complete fact-level rights review
code
Reviewed qualifying license recorded — check scope and conditions
data
Unknown — no complete fact-level rights review
documentation
Unknown — no complete fact-level rights review

No complete system-rights review is recorded for this release.

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.2. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.2): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.2. Unknown means unassessed or insufficient evidence.
Public materials checklist for NVIDIA parakeet-unified-en-0.6b
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicThe ungated repository holds the .nemo checkpoint.2
Inference codeIs code for running the model published?PublicThe card documents offline and streaming inference with NVIDIA NeMo, including a chunked RNN-T streaming script and a pipeline configuration in the Apache 2.0 NeMo Speech repository.15
Training codeIs the code used to train the model published?UnknownSources disagree. The card says only inference is supported for now and that the unified training pipeline will be released later, while the accompanying paper's abstract says the Unified ASR framework is open-sourced. NeMo Speech 3.0 release notes mention Unified RNN-T streaming inference. Training-code availability was not confirmed.137
Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness.PartialThe card states about 250,000 hours of audio, mostly from the English portion of Granary (YouTube-Commons, YODAS2, MOSEL, and LibriLight, with hours per source), plus a list of other corpora, and gives collection and labeling methods (human-recorded audio; a mix of human and ASR-generated transcripts). This review confirms disclosure, not completeness.18
Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable.PartialGranary is published as CC-BY-4.0 transcript manifests that point to separately downloaded source audio; access terms for the other listed corpora were not reviewed.81
Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient.UnknownNot assessed.
Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above.UnknownNot assessed.
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialThe card describes the architecture, joint offline and streaming training with chunked attention masks, dynamic chunked convolutions, and the consistency loss; hyperparameters and schedules are not given on the card.13
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PartialThe card publishes offline and streaming word error rates on Open ASR Leaderboard datasets and names the text normalizer version used; evaluation scripts were not verified.1

What it is useful for

Transcribing English audio in batch or streaming mode, for example in voice assistants, live captioning, and conversational AI. For 80 ms latency the card recommends a different model, nemotron-speech-streaming-en-0.6b.1Fact reviewed Oct 1, 2026

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The card lists NeMo 2.7.3 as the runtime, Linux as the supported operating system, NVIDIA Volta, Ampere, Hopper, and Blackwell GPUs as compatible (tested on V100, A100, A6000, and DGX Spark), and recommended left, chunk, and right context settings for each streaming latency.1

Organization context

Provenance and derivatives

Trained by NVIDIA. Most training data comes from the English portion of Granary, whose card says transcripts were pseudo-labeled with Whisper-large-v3 and that punctuation and capitalization were restored with Qwen-2.5-7B, both third-party models. No third-party model weights are documented as part of this checkpoint.18

Other releases in the NVIDIA Parakeet family

NVIDIA Parakeet family overview

What this catalog does not know

Unknown means the sources reviewed for this record do not document it. It is not evidence that something does not exist.
  • Training code: unknown.
  • Complete training pipeline: unknown.
  • Legacy data assessment (v0.1): unknown.

Have a primary source? How to report a correction.

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

Published by NVIDIA in its Hugging Face organization under the NVIDIA Open Model License and announced in NVIDIA's NeMo Speech repository. NVIDIA's principal executive offices are in Santa Clara, California, per its Form 10-Q for the quarter ended July 26, 2026. Third-party models used to label training data are recorded under provenance.169

Assessed Oct 1, 2026

Sources

  1. 1.
    nvidia/parakeet-unified-en-0.6b model card (external site: huggingface.co)

    NVIDIA (Hugging Face) · Model card · published Apr 7, 2026 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  2. 2.
    Hugging Face model metadata for nvidia/parakeet-unified-en-0.6b (external site: huggingface.co)

    Hugging Face · Other · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  3. 3.
    Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization (external site: arxiv.org)

    NVIDIA (arXiv) · Paper · published Apr 21, 2026 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  4. 4.
    NVIDIA Open Model License Agreement (external site: nvidia.com)

    NVIDIA · License · published Oct 24, 2025 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  5. 5.
    NVIDIA-NeMo/Speech LICENSE (Apache License 2.0) (external site: raw.githubusercontent.com)

    NVIDIA (GitHub) · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  6. 6.
    NVIDIA-NeMo/Speech README (external site: raw.githubusercontent.com)

    NVIDIA (GitHub) · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  7. 7.
    NVIDIA-NeMo/Speech release v3.0.0 (external site: github.com)

    NVIDIA (GitHub) · Release notes · published Aug 7, 2026 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  8. 8.
    nvidia/Granary dataset card (external site: huggingface.co)

    NVIDIA (Hugging Face) · Dataset card · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  9. 9.
    NVIDIA Corporation Form 10-Q for the quarter ended July 26, 2026 (external site: sec.gov)

    U.S. Securities and Exchange Commission (EDGAR) · Filing · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project