Hub
Training data and datasets
Open pretraining datasets in the catalog and the organizations that steward them, how many of them build on Common Crawl, and why a dataset's license and documentation are not permission to reuse what it contains.
Many of the open text datasets in this catalog start from the same source. Common Crawl, a 501(c)(3) nonprofit, crawls the web and freely provides an archive of crawl data collected since 2008. Hugging Face's FineWeb is built by filtering and deduplicating Common Crawl snapshots. Together AI's RedPajama-V2 contains documents from 84 Common Crawl snapshots, with quality signals for part of the corpus and the IDs of duplicate documents so that you can filter it yourself. Essential AI's Essential-Web, also built from Common Crawl, labels every document with subject, page type, complexity, and quality metadata. Ai2's Dolma mixes web content with academic publications, code, books, and encyclopedic material.14675
A dataset's license covers the compilation, not every work inside it. FineWeb and Dolma are released under the Open Data Commons Attribution License, but FineWeb's card says its use is also subject to Common Crawl's Terms of Use, and Dolma's card says you are also bound by the licenses and terms of the original data sources. RedPajama-V2 points to Common Crawl's terms for its data and uses Apache 2.0 only for its code. Common Crawl's terms in turn say you must evaluate and bear the risks of using crawled content, recommend legal advice before any use, including commercial use, and require you to respect the copyrights and other rights of third parties in that material.4562
Some datasets come with use restrictions instead. Downloading Nemotron-CC-v2 from Hugging Face requires accepting the NVIDIA Data Agreement for Model Training and confirming that the data will be used only for model training. The agreement makes the data available solely for internal training of your own AI models, bars redistributing it, and says that NVIDIA does not grant any rights to copyrighted material the datasets may contain.89
Documentation does not guarantee lasting availability. EleutherAI's Pile was described in a paper and a datasheet, and its card says the license depends on which of its component datasets you use. A 2025 paper describing the Common Pile, a dataset of public domain and openly licensed text, says that the use of unlicensed training data has previously resulted in DMCA takedowns of datasets such as the Pile. Terms can also change: Ai2 switched Dolma to the ODC-By license in April 2024.10115
Web data includes personal information, and stewards document different responses. FineWeb replaces email addresses and public IP addresses, says some personal information is still likely to remain, and offers a removal form; Dolma's card also links a form for requesting removal of personal data. Website owners who do not want Common Crawl to crawl their sites can block its CCBot crawler in robots.txt.453
What this hub covers
Covers open pretraining and research datasets in the catalog, the organizations that collect or maintain them, and the documents that state their sources, processing, licenses, and access conditions. It does not decide whether any use of a dataset is lawful and is not legal advice. Featured records are examples chosen to cover the subject, not a ranking or a complete list.
Reading path
- Training-data disclosuresStart here. Separates which data was used, where it came from, access, and terms.
- License scopeWhy a dataset license may not cover the works collected in it.
- Availability is not permissionHow the catalog keeps what you can download separate from what you may do with it.
- Gated downloadWhat it means when a dataset asks you to sign in and accept terms first.
- What open weight and open source meanHow data terms fit alongside the terms for weights and code in one release.
- UnknownA gap in a disclosure means not assessed or not documented, never "no".
- Common Crawl corpus recordThe catalog record for the web archive that many of these datasets build on.
In the catalog
- Datasets in the catalogOpen dataset records, each with its sources, licenses, and access conditions.
- Organizations working with dataOrganizations whose records list data as a sector.
- Data servicesOrganizations that collect, label, or supply data for AI.
Featured records
Examples chosen to cover the subject; not a ranking or a complete list.
- Common CrawlOrganization
- Ai2 (Allen Institute for AI)Organization
- EleutherAIOrganization
- Hugging FaceOrganization
- Together AIOrganization
- Common Crawl corpusOpen model or tool
- FineWebOpen model or tool
- DolmaOpen model or tool
- RedPajamaOpen model or tool
- Essential-Web v1.0Open model or tool
- Nemotron-CCOpen model or tool
- The PileOpen model or tool
Primary documents
- Terms of Use (external site: commoncrawl.org)Common Crawl FoundationCommon Crawl's Terms of Use, which put the risk and the duty to respect third-party rights on whoever uses the crawled content.
- FineWeb dataset card (external site: huggingface.co)Hugging Face (Hugging Face Hub)A dataset card that states its own license, the upstream terms that also apply, and its handling of personal information.
- Dolma dataset card (external site: huggingface.co)Ai2 (Hugging Face)Dolma's card, including its license change and the statement that the original sources' terms still apply.
- RedPajama-Data-V2 dataset card (external site: huggingface.co)Together AI (Hugging Face)Shows a dataset released with quality signals and duplicate IDs rather than a single filtered version, with separate terms for data and code.
- NVIDIA Data Agreement for Model Training (external site: huggingface.co)NVIDIA (Hugging Face)An example of a dataset agreement that limits use to model training and disclaims any grant of rights to copyrighted content.
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text (arXiv 2506.05209) (external site: arxiv.org)arXivExplains what "openly licensed" text means and why unlicensed training data limits sharing datasets.
- CCBot (external site: commoncrawl.org)Common Crawl FoundationHow website owners can identify Common Crawl's crawler and block it in robots.txt.
Sources · reviewed Oct 2, 2026
Support Us
Help keep USASI useful.
Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).
About supporting this project