USASI
DatasetDataset

FineWeb

Version v1.4.0

Maintained by Hugging Face (FineData, HuggingFaceFW)15

FineWeb is Hugging Face's English web-text pretraining dataset, built by extracting, filtering, and deduplicating pages from Common Crawl snapshots dating back to 2013. The card describes it as more than 18.5 trillion tokens (GPT-2 tokenizer), up from about 15 trillion at first release in April 2024; version 1.4.0 (July 2025) added the Common Crawl snapshots from January to June 2025. Hugging Face also published FineWeb-Edu, an educational subset.12

Last reviewedEntry updated Documented release Apr 21, 2024

Availability and license

Overall availability

Public

Downloadable from the Hugging Face Hub without gating (via datasets, huggingface_hub, or datatrove). Use is also subject to Common Crawl's Terms of Use, per the card.1

Availability is separate from permission: read the license before using or redistributing.

The card releases the dataset under ODC-By 1.0 and states that its use is also subject to Common Crawl's Terms of Use. It notes that some domains were removed in version 1.3.0 in response to a cease-and-desist notice.1

Public materials checklist

Items for a dataset under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for FineWeb
ItemStatusNotes and evidence
AccessCan the data be obtained, and on what terms?PublicThe full dataset, per-snapshot subsets, and sample subsets are downloadable without gating; previous versions remain available on named branches.1
ProvenanceAre the data's origins documented?PublicEvery record keeps its Common Crawl snapshot, source URL, crawl date, and WARC file path, and the card lists the processing pipeline step by step.1
DocumentationIs there a datasheet, card, or equivalent documentation?PublicThe dataset card, a NeurIPS 2024 Datasets and Benchmarks paper, and a runnable datatrove script that reproduces the processing pipeline are published; ablation models and evaluation results are also released.123
LicensingAre the licensing terms stated?PublicODC-By 1.0 for the data, also subject to Common Crawl's Terms of Use.1
Stated limitationsDoes the documentation state known limitations or risks?PublicThe card states that despite URL filtering the data likely still contains toxic or harmful content and personal information, that email and public IP addresses were anonymized, and that code content is likely not prevalent.1

What it is useful for

The card describes it as a research artifact for pretraining large language models on public web data; each Common Crawl snapshot can be loaded separately, and 10B, 100B, and 350B-token random samples are provided.1

Organization context

Provenance and derivatives

Derived from Common Crawl web crawls. Hugging Face extracted text from the WARC files with Trafilatura, applied URL, English-language (fastText), and quality filters (Gopher, C4, and FineWeb-specific heuristics), deduplicated each snapshot separately with MinHash, and anonymized email and public IP addresses.13

  • Derived from: Common Crawl — Source web crawls (snapshots from 2013 onward).

Catalog records that name this entry in their provenance:

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

FineWeb is created and published by Hugging Face's HuggingFaceFW organization ("FineData"), which describes itself as part of the Hugging Face Science team, and it is processed with Hugging Face's datatrove library. Hugging Face's terms of service identify Hugging Face, Inc., a Delaware corporation, as the provider of its services, and its privacy policy states that the company is located in the United States; see the hugging-face organization record for the dual-country assessment. The underlying web data comes from the Common Crawl Foundation's crawls and is recorded under provenance.1567

Assessed Sep 29, 2026

Sources

  1. 1.
    HuggingFaceFW/fineweb dataset card (external site: huggingface.co)

    Hugging Face · Dataset card · accessed Sep 29, 2026

  2. 2.
    The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale (arXiv 2406.17557) (external site: arxiv.org)

    Hugging Face (arXiv) · Paper · published Jun 25, 2024 · accessed Sep 29, 2026

  3. 3.
    datatrove examples/fineweb.py (external site: raw.githubusercontent.com)

    Hugging Face · Repository · accessed Sep 29, 2026

  4. 4.
    huggingface/datatrove LICENSE (external site: raw.githubusercontent.com)

    Hugging Face · License · accessed Sep 29, 2026

  5. 5.
    HuggingFaceFW (FineData) (external site: huggingface.co)

    Hugging Face · Official page · accessed Sep 29, 2026

  6. 6.
    Terms of Service (external site: huggingface.co)

    Hugging Face · Official page · accessed Sep 29, 2026

  7. 7.
    Hugging Face Privacy Policy (external site: huggingface.co)

    Hugging Face · Official page · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project