FineWeb
Version v1.4.0
Maintained by Hugging Face (FineData, HuggingFaceFW)15
FineWeb is Hugging Face's English web-text pretraining dataset, built by extracting, filtering, and deduplicating pages from Common Crawl snapshots dating back to 2013. The card describes it as more than 18.5 trillion tokens (GPT-2 tokenizer), up from about 15 trillion at first release in April 2024; version 1.4.0 (July 2025) added the Common Crawl snapshots from January to June 2025. Hugging Face also published FineWeb-Edu, an educational subset.12
- Dataset hub: Dataset card (Hugging Face) (external site: huggingface.co)
- Paper: FineWeb paper (arXiv 2406.17557) (external site: arxiv.org)
- Repository: Processing script (datatrove examples/fineweb.py) (external site: github.com)
- License: ODC-By 1.0 license (external site: opendatacommons.org)
Availability and license
Overall availability
Downloadable from the Hugging Face Hub without gating (via datasets, huggingface_hub, or datatrove). Use is also subject to Common Crawl's Terms of Use, per the card.1
Availability is separate from permission: read the license before using or redistributing.
Open Data Commons Attribution License v1.0 (external site: opendatacommons.org)1
Apache License 2.0 (datatrove processing library) (external site: raw.githubusercontent.com)4
The card releases the dataset under ODC-By 1.0 and states that its use is also subject to Common Crawl's Terms of Use. It notes that some domains were removed in version 1.3.0 in response to a cease-and-desist notice.1
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| AccessCan the data be obtained, and on what terms? | Public | The full dataset, per-snapshot subsets, and sample subsets are downloadable without gating; previous versions remain available on named branches.1 |
| ProvenanceAre the data's origins documented? | Public | Every record keeps its Common Crawl snapshot, source URL, crawl date, and WARC file path, and the card lists the processing pipeline step by step.1 |
| DocumentationIs there a datasheet, card, or equivalent documentation? | Public | The dataset card, a NeurIPS 2024 Datasets and Benchmarks paper, and a runnable datatrove script that reproduces the processing pipeline are published; ablation models and evaluation results are also released.123 |
| LicensingAre the licensing terms stated? | Public | ODC-By 1.0 for the data, also subject to Common Crawl's Terms of Use.1 |
| Stated limitationsDoes the documentation state known limitations or risks? | Public | The card states that despite URL filtering the data likely still contains toxic or harmful content and personal information, that email and public IP addresses were anonymized, and that code content is likely not prevalent.1 |
What it is useful for
The card describes it as a research artifact for pretraining large language models on public web data; each Common Crawl snapshot can be loaded separately, and 10B, 100B, and 350B-token random samples are provided.1
Organization context
Provenance and derivatives
Derived from Common Crawl web crawls. Hugging Face extracted text from the WARC files with Trafilatura, applied URL, English-language (fastText), and quality filters (Gopher, C4, and FineWeb-specific heuristics), deduplicated each snapshot separately with MinHash, and anonymized email and public IP addresses.13
- Derived from: Common Crawl — Source web crawls (snapshots from 2013 onward).
Catalog records that name this entry in their provenance:
U.S. eligibility
Eligible · basis: U.S.-governed project
FineWeb is created and published by Hugging Face's HuggingFaceFW organization ("FineData"), which describes itself as part of the Hugging Face Science team, and it is processed with Hugging Face's datatrove library. Hugging Face's terms of service identify Hugging Face, Inc., a Delaware corporation, as the provider of its services, and its privacy policy states that the company is located in the United States; see the hugging-face organization record for the dual-country assessment. The underlying web data comes from the Common Crawl Foundation's crawls and is recorded under provenance.1567
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.