Essential-Web v1.0
Version v1.0
Maintained by Essential AI12
Essential-Web v1.0 is a web text dataset of about 24 trillion tokens in 23.6 billion documents, built by Essential AI from 101 Common Crawl snapshots. Every document carries metadata from a 12-category taxonomy covering subject (Free Decimal Correspondence, a Dewey Decimal-inspired scheme), page type, reasoning depth, education level, and quality, so subsets can be selected with metadata filters.12
- Dataset hub: Dataset card (Hugging Face) (external site: huggingface.co)
- Paper: Essential-Web v1.0 paper (arXiv 2506.14111) (external site: arxiv.org)
- Model hub: EAI-Distill-0.5b labeling model (external site: huggingface.co)
- Repository: eai-taxonomy code (external site: github.com)
Availability and license
Overall availability
Downloadable from Hugging Face without gating. The card asks users to also follow the Common Crawl terms of use.1
Availability is separate from permission: read the license before using or redistributing.
The card says Essential AI's contributions are released under ODC-By, that users should also follow the Common Crawl terms of use, and that Essential AI does not alter the license of any underlying data.1
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| AccessCan the data be obtained, and on what terms? | Public | 1 |
| ProvenanceAre the data's origins documented? | Public | The card lists the Common Crawl snapshot ranges and the processing steps: global and MinHash deduplication, quality signals from a RedPajama-Data-V2-style pipeline and the DCLM-baseline fastText classifier, filtering, and taxonomy labeling.1 |
| DocumentationIs there a datasheet, card, or equivalent documentation? | Public | The card documents the schema and taxonomy codes; the paper describes the method.12 |
| LicensingAre the licensing terms stated? | Public | ODC-By for Essential AI's contributions, alongside Common Crawl's terms of use.1 |
| Stated limitationsDoes the documentation state known limitations or risks? | Unknown | The paper has no dedicated limitations section; it notes that its small labeling model does worse than its teacher on some categories, such as extraction artifacts and education level.2 |
What it is useful for
Building domain-specific pretraining sets without custom classifiers by filtering on the taxonomy labels. Essential AI has published math, code, STEM, and medical subsets curated this way.1
Organization context
Provenance and derivatives
Built from Common Crawl: 89 snapshots from the DCLM pool (CC-MAIN-2013-20 to CC-MAIN-2022-49) plus 12 later snapshots (CC-MAIN-2023-06 to CC-MAIN-2024-38), all extracted with resiliparse. Taxonomy labels were produced by EAI-Distill-0.5b, which Essential AI fine-tuned from Qwen2.5-0.5B-Instruct (from Alibaba Cloud's Qwen family) on labels generated by Qwen2.5-32B-Instruct; the dataset card refers to the labeler as EAI-Taxonomy-0.5b.1235
- Derived from: Common Crawl — Source of all documents (101 snapshots).
- Derived from: EAI-Distill-0.5b (external site: huggingface.co) — Essential AI labeling model fine-tuned from Qwen2.5-0.5B-Instruct.
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.