Independent project. Not a U.S. government website.

USASI

Hub

Training data and datasets

Open pretraining datasets in the catalog and the organizations that steward them, how many of them build on Common Crawl, and why a dataset's license and documentation are not permission to reuse what it contains.

Many of the open text datasets in this catalog start from the same source. Common Crawl, a 501(c)(3) nonprofit, crawls the web and freely provides an archive of crawl data collected since 2008. Hugging Face's FineWeb is built by filtering and deduplicating Common Crawl snapshots. Together AI's RedPajama-V2 contains documents from 84 Common Crawl snapshots, with quality signals for part of the corpus and the IDs of duplicate documents so that you can filter it yourself. Essential AI's Essential-Web, also built from Common Crawl, labels every document with subject, page type, complexity, and quality metadata. Ai2's Dolma mixes web content with academic publications, code, books, and encyclopedic material.14675

A dataset's license covers the compilation, not every work inside it. FineWeb and Dolma are released under the Open Data Commons Attribution License, but FineWeb's card says its use is also subject to Common Crawl's Terms of Use, and Dolma's card says you are also bound by the licenses and terms of the original data sources. RedPajama-V2 points to Common Crawl's terms for its data and uses Apache 2.0 only for its code. Common Crawl's terms in turn say you must evaluate and bear the risks of using crawled content, recommend legal advice before any use, including commercial use, and require you to respect the copyrights and other rights of third parties in that material.4562

Some datasets come with use restrictions instead. Downloading Nemotron-CC-v2 from Hugging Face requires accepting the NVIDIA Data Agreement for Model Training and confirming that the data will be used only for model training. The agreement makes the data available solely for internal training of your own AI models, bars redistributing it, and says that NVIDIA does not grant any rights to copyrighted material the datasets may contain.89

Documentation does not guarantee lasting availability. EleutherAI's Pile was described in a paper and a datasheet, and its card says the license depends on which of its component datasets you use. A 2025 paper describing the Common Pile, a dataset of public domain and openly licensed text, says that the use of unlicensed training data has previously resulted in DMCA takedowns of datasets such as the Pile. Terms can also change: Ai2 switched Dolma to the ODC-By license in April 2024.10115

Web data includes personal information, and stewards document different responses. FineWeb replaces email addresses and public IP addresses, says some personal information is still likely to remain, and offers a removal form; Dolma's card also links a form for requesting removal of personal data. Website owners who do not want Common Crawl to crawl their sites can block its CCBot crawler in robots.txt.453

What this hub covers

Covers open pretraining and research datasets in the catalog, the organizations that collect or maintain them, and the documents that state their sources, processing, licenses, and access conditions. It does not decide whether any use of a dataset is lawful and is not legal advice. Featured records are examples chosen to cover the subject, not a ranking or a complete list.

Reading path

  1. 1.Training-data disclosuresStart here. Separates which data was used, where it came from, access, and terms.
  2. 2.License scopeWhy a dataset license may not cover the works collected in it.
  3. 3.Availability is not permissionHow the catalog keeps what you can download separate from what you may do with it.
  4. 4.Gated downloadWhat it means when a dataset asks you to sign in and accept terms first.
  5. 5.What open weight and open source meanHow data terms fit alongside the terms for weights and code in one release.
  6. 6.UnknownA gap in a disclosure means not assessed or not documented, never "no".
  7. 7.Common Crawl corpus recordThe catalog record for the web archive that many of these datasets build on.

In the catalog

Examples chosen to cover the subject; not a ranking or a complete list.

Primary documents

Sources · reviewed Oct 2, 2026

  1. 1.
    About Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Official page · accessed Oct 2, 2026

  2. 2.
    Terms of Use (external site: commoncrawl.org)

    Common Crawl Foundation · License · accessed Oct 2, 2026

  3. 3.
    CCBot (external site: commoncrawl.org)

    Common Crawl Foundation · Official page · accessed Oct 2, 2026

  4. 4.
    FineWeb dataset card (external site: huggingface.co)

    Hugging Face (Hugging Face Hub) · Dataset card · accessed Oct 2, 2026

  5. 5.
    Dolma dataset card (external site: huggingface.co)

    Ai2 (Hugging Face) · Dataset card · accessed Oct 2, 2026

  6. 6.
    RedPajama-Data-V2 dataset card (external site: huggingface.co)

    Together AI (Hugging Face) · Dataset card · accessed Oct 2, 2026

  7. 7.
    Essential-Web v1.0 dataset card (external site: huggingface.co)

    Essential AI (Hugging Face) · Dataset card · accessed Oct 2, 2026

  8. 8.
  9. 9.
    NVIDIA Data Agreement for Model Training (external site: huggingface.co)

    NVIDIA (Hugging Face) · License · accessed Oct 2, 2026

  10. 10.
    Dataset Card for The Pile (external site: huggingface.co)

    EleutherAI (Hugging Face) · Dataset card · accessed Oct 2, 2026

  11. 11.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project