Independent project. Not a U.S. government website.

USASI
Model release

NVIDIA parakeet-tdt-0.6b-v3

Release in the NVIDIA Parakeet family · version parakeet-tdt-0.6b-v3

Maintained by NVIDIA (NeMo)16

parakeet-tdt-0.6b-v3 is a 600-million-parameter multilingual speech recognition model with a FastConformer encoder and a TDT decoder, which predicts tokens and their durations jointly. It extends parakeet-tdt-0.6b-v2 from English to 25 European languages, detects the spoken language automatically, and outputs punctuated, capitalized text with word- and segment-level timestamps. NVIDIA released it on Hugging Face on August 14, 2025.13Fact reviewed Oct 1, 2026

Last reviewedEntry updated Documented release Aug 14, 2025

Availability and license

Overall availability

Public

Downloadable from Hugging Face without an access gate; the card states the model is ready for commercial and non-commercial use.21

Availability fact review: Oct 1, 2026

Availability is separate from permission: read the license before using or redistributing.

The card's governing terms put the model under CC-BY-4.0, which grants a worldwide, royalty-free license to reproduce, share, and adapt the material and requires attribution and license notices when sharing it; the card also states the model is ready for commercial and non-commercial use. The NeMo toolkit used to run and train it is a separate Apache 2.0 repository.145Fact reviewed Oct 1, 2026

Component reuse rights

A readable or downloadable component is not automatically reusable. These indicators concern recorded license evidence, not system certification.
weights
Terms or license combinations require review
code
Reviewed qualifying license recorded — check scope and conditions
data
Unknown — no complete fact-level rights review
documentation
Unknown — no complete fact-level rights review

No complete system-rights review is recorded for this release.

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.2. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.2): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.2. Unknown means unassessed or insufficient evidence.
Public materials checklist for NVIDIA parakeet-tdt-0.6b-v3
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicThe ungated repository holds safetensors and .nemo checkpoints plus a q8_0 GGUF file.2
Inference codeIs code for running the model published?PublicThe card documents inference with NVIDIA NeMo, the NeMo-Speech.cpp C++ runtime, and Hugging Face Transformers, including timestamps, long-form, and streaming examples.15
Training codeIs the code used to train the model published?PublicThe card links the NeMo transducer training script and FastConformer hybrid TDT-CTC configuration it was trained with; both are in the Apache 2.0 NeMo Speech repository.1785
Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness.PartialThe card describes about 660,000 hours of pseudo-labeled Granary data (from YTC, MOSEL, and YODAS) and about 10,000 hours of human-transcribed NeMo ASR Set 3.0 data, listing the component corpora; the Granary card and paper describe the labeling pipeline. This review confirms disclosure, not completeness across every provenance, labeling, and processing requirement.1910
Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable.PartialGranary is published as CC-BY-4.0 transcript manifests that point to separately downloaded source audio; NeMo ASR Set 3.0 is described as an in-house set, and access terms for its component corpora were not reviewed.91
Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient.UnknownNot assessed.
Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above.UnknownNot assessed.
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialThe card gives the initialization (an NVIDIA CTC multilingual checkpoint pretrained on Granary), step counts and GPU counts for both stages, temperature-based data balancing, and the tokenizer setup; the recipe for the initial checkpoint is not described there.1
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PartialThe card publishes word error rates on FLEURS, MLS, CoVoST, and Open ASR Leaderboard datasets; evaluation scripts were not verified.1

What it is useful for

The card lists speech-to-text uses such as conversational AI, voice assistants, transcription services, subtitle generation, and voice analytics, and documents long-form transcription with a local-attention mode (up to three hours of audio).1Fact reviewed Oct 1, 2026

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The card lists NeMo 2.4 as the runtime, Linux as the supported operating system, and NVIDIA Volta, Ampere, Hopper, and Blackwell GPUs as compatible, and shows a local NeMo-Speech.cpp workflow using the q8_0-quantized GGUF file.1

Organization context

Provenance and derivatives

Trained by NVIDIA from an NVIDIA CTC multilingual checkpoint pretrained on Granary. Granary was built by the NVIDIA NeMo team with CMU and FBK teams; its card says transcripts were pseudo-labeled with Whisper-large-v3 and that punctuation and capitalization were restored with Qwen-2.5-7B, both third-party models. No third-party model weights are documented as part of this checkpoint.19

Other releases in the NVIDIA Parakeet family

NVIDIA Parakeet family overview

What this catalog does not know

Unknown means the sources reviewed for this record do not document it. It is not evidence that something does not exist.
  • Complete training pipeline: unknown.
  • Legacy data assessment (v0.1): unknown.

Have a primary source? How to report a correction.

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

Published by NVIDIA in its Hugging Face organization, trained with NVIDIA's NeMo toolkit per the card, and announced in NVIDIA's NeMo Speech repository. NVIDIA's principal executive offices are in Santa Clara, California, per its Form 10-Q for the quarter ended July 26, 2026. Third-party models used to label training data are recorded under provenance.1611

Assessed Oct 1, 2026

Sources

  1. 1.
    nvidia/parakeet-tdt-0.6b-v3 model card (external site: huggingface.co)

    NVIDIA (Hugging Face) · Model card · published Aug 14, 2025 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  2. 2.
    Hugging Face model metadata for nvidia/parakeet-tdt-0.6b-v3 (external site: huggingface.co)

    Hugging Face · Other · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  3. 3.
    Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST (external site: arxiv.org)

    NVIDIA (arXiv) · Paper · published Sep 17, 2025 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  4. 4.
    Creative Commons Attribution 4.0 International Public License (legal code) (external site: creativecommons.org)

    Creative Commons · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  5. 5.
    NVIDIA-NeMo/Speech LICENSE (Apache License 2.0) (external site: raw.githubusercontent.com)

    NVIDIA (GitHub) · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  6. 6.
    NVIDIA-NeMo/Speech README (external site: raw.githubusercontent.com)

    NVIDIA (GitHub) · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  7. 7.
    NeMo Speech examples/asr/asr_transducer/speech_to_text_rnnt_bpe.py (external site: raw.githubusercontent.com)

    NVIDIA (GitHub) · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  8. 8.
    NeMo Speech fastconformer_hybrid_tdt_ctc_bpe.yaml configuration (external site: raw.githubusercontent.com)

    NVIDIA (GitHub) · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  9. 9.
    nvidia/Granary dataset card (external site: huggingface.co)

    NVIDIA (Hugging Face) · Dataset card · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  10. 10.
    Granary: Speech Recognition and Translation Dataset in 25 European Languages (external site: arxiv.org)

    NVIDIA (arXiv) · Paper · published May 19, 2025 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  11. 11.
    NVIDIA Corporation Form 10-Q for the quarter ended July 26, 2026 (external site: sec.gov)

    U.S. Securities and Exchange Commission (EDGAR) · Filing · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project