Nemotron-CC
Project record
Nemotron-CC is NVIDIA's English pretraining dataset built from Common Crawl. The original release (December 2024) has 6.3T tokens, 4.4T globally deduplicated original tokens and 1.9T synthetic tokens, drawn from 99 Common Crawl snapshots (CC-MAIN-2013-20 through CC-MAIN-2024-30) and split into five quality buckets. Nemotron-CC-v2 (August 2025) and Nemotron-CC-v2.1 (December 2025), published on Hugging Face, add newer snapshots, further synthetic rephrasings, translated question-answer data, and other subsets.13245
- Dataset hub: Nemotron-CC (original release, hosted by Common Crawl) (external site: data.commoncrawl.org)
- Dataset hub: Nemotron-CC-v2 (Hugging Face) (external site: huggingface.co)
- Dataset hub: Nemotron-CC-v2.1 (Hugging Face) (external site: huggingface.co)
- Paper: Paper: Nemotron-CC (arXiv 2412.02595) (external site: arxiv.org)
- Website: NVIDIA research page (external site: research.nvidia.com)
- License: NVIDIA Data Agreement for Model Training (external site: huggingface.co)
Availability and license
Overall availability
The original release is downloadable from Common Crawl's data site as partitioned .jsonl.zstd files (about 10.4 TiB). Nemotron-CC-v2 and v2.1 are gated on Hugging Face: users must share contact information (including company and institutional email) and agree to the NVIDIA Data Agreement for Model Training, confirming that they intend to use the data for model training only. The Hugging Face API reports both gates as manual, meaning the dataset authors approve each access request.145678
Availability is separate from permission: read the license before using or redistributing.
Common Crawl Terms of Use (original Nemotron-CC release) (external site: commoncrawl.org)310
NVIDIA Data Agreement for Model Training (Nemotron-CC-v2 and v2.1) (external site: huggingface.co)459
The paper states the original release is under the Common Crawl Terms of Use, which grant a limited, non-transferable license and note that crawled content may carry its owners' terms. The NVIDIA Data Agreement for Model Training (version of August 15, 2025) makes the data available solely for internal training of the user's AI models; it prohibits redistributing or sublicensing the datasets, grants no rights in copyrighted material contained in them, lets either party terminate on 30 days' notice (after which copies must be deleted), and is governed by U.S. and Delaware law. The v2 card lists the models used to generate synthetic data across its dataset collection and states that models trained on the data may be subject to the Qwen and DeepSeek license agreements; the v2.1 card also names the Phi-4 license agreement.310945
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| AccessCan the data be obtained, and on what terms? | Partial | Original release downloadable from Common Crawl. v2 and v2.1 require sharing contact information and accepting NVIDIA's agreement on Hugging Face, and the Hub reports manual approval of access requests for both.145678 |
| ProvenanceAre the data's origins documented? | Public | The Common Crawl page lists the crawls included and the quality and synthetic-type partitions; each record keeps its source URL and WARC record ID. The v2 and v2.1 cards name the added snapshots and the models used to generate synthetic data.145 |
| DocumentationIs there a datasheet, card, or equivalent documentation? | Public | The construction pipeline (extraction, classifier ensembling, quality bucketing, synthetic rephrasing, with prompt templates in an appendix) is described in the paper; v2 and v2.1 have dataset cards.345 |
| LicensingAre the licensing terms stated? | Public | Terms documented for each release; neither set of terms is an open data license.3109 |
| Stated limitationsDoes the documentation state known limitations or risks? | Public | The paper's Limitations section says the rephrased data was not checked for factual accuracy or fidelity, the methods were tried only on English text, not every pipeline step was ablated, and the dataset was not decontaminated against benchmarks.3 |
What it is useful for
Organization context
Provenance and derivatives
Built by NVIDIA from Common Crawl web archives. For the original release, synthetic rephrasings were generated with the instruct version of Mistral NeMo 12B. The v2 card's per-dataset table lists Mistral-Nemo-12B-Instruct and Qwen3-30B-A3B for Nemotron-CC-v2. The v2.1 card lists Qwen3-30B-A3B for its synthetic and translated subsets, and says its STEM question-answer subset was built from documents in Essential-Web.3145
- Derived from: Common Crawl — Source web archive; the original release covers crawls CC-MAIN-2013-20 through CC-MAIN-2024-30.
U.S. eligibility
Eligible · basis: U.S.-governed project
The dataset is created and published by NVIDIA: the paper's authors are listed with an NVIDIA affiliation, the v2 card names NVIDIA as data developer, and the v2.1 card names NVIDIA Corporation as dataset owner. NVIDIA Corporation's principal executive offices are in Santa Clara, California, per its Form 10-Q for the quarter ended July 26, 2026.34511
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.