Dolma
Project record
Maintained by Ai2 (Allen Institute for AI)1112
Dolma is Ai2's family of English pretraining corpora built from web pages, academic publications, code, and encyclopedic text, used to train the OLMo models. The first release (v1, August 2023) was updated through v1.7 (April 2024); Dolma 3, used for Olmo 3, consists of a pool of about 9.3 trillion tokens and curated pretraining, mid-training, and long-context mixes. A Dolma 3.5 pool with additional sources and filtering has since been published.129115
- Dataset hub: Dolma v1.x (Hugging Face) (external site: huggingface.co)
- Dataset hub: Dolma 3 pool (Hugging Face) (external site: huggingface.co)
- Paper: Dolma paper and datasheet (arXiv 2402.00159) (external site: arxiv.org)
- Repository: Dolma toolkit (external site: github.com)
- Repository: Dolma 3 construction repository (external site: github.com)
- Documentation: Dolma documentation (external site: docs.allenai.org)
Availability and license
Overall availability
Downloadable from the Hugging Face Hub without gating. The v1.x card states that users are also bound by the license terms of the original data sources, and the Dolma 3 pool repository hosts only its Common Crawl and olmOCR PDF portions, linking the other sources to their original repositories.12
Availability is separate from permission: read the license before using or redistributing.
Open Data Commons Attribution License v1.0 (external site: huggingface.co)1235
Apache License 2.0 (Dolma toolkit and Dolma 3 construction code) (external site: raw.githubusercontent.com)810
Dolma moved to ODC-BY in April 2024, per the v1.x dataset card, which also states that users remain bound by the terms of the original data sources. The Dolma 3 cards state the data is intended for research and educational use under Ai2's Responsible Use Guidelines.12
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| AccessCan the data be obtained, and on what terms? | Public | All versions are downloadable from Hugging Face without gating; some Dolma 3 pool sources are obtained from their original repositories.12 |
| ProvenanceAre the data's origins documented? | Public | Dataset cards list sources (for example Common Crawl, code, scientific PDFs, arXiv, math web pages, and Wikipedia) with token and document counts per source.124 |
| DocumentationIs there a datasheet, card, or equivalent documentation? | Public | The Dolma paper includes a datasheet appendix; the dolma3 repository documents how the Dolma 3 datasets were built.619 |
| LicensingAre the licensing terms stated? | Public | ODC-BY for the data; Apache 2.0 for the toolkit code.128 |
| Stated limitationsDoes the documentation state known limitations or risks? | Public | The paper's Limitations section notes the English-only focus, incomplete coverage of curation practices, ablations run only at 1B scale, and that full manual inspection of the corpus is not feasible. The Olmo 3 7B mix card notes that some PDFs were redacted after training.64 |
What it is useful for
Organization context
Provenance and derivatives
Assembled by Ai2 from existing sources. Dolma v1.7 draws on Common Crawl, RefinedWeb, StarCoder, C4, Reddit, peS2o, arXiv and StackExchange via RedPajama, and Wikipedia, among others. The Dolma 3 pool combines Common Crawl, olmOCR-processed science PDFs, Stack-Edu, arXiv, FineMath, and Wikipedia.12
- Derived from: Common Crawl — Main web source across versions.
- Derived from: Stack-Edu (external site: huggingface.co) — Code source in the Dolma 3 pool.
- Derived from: FineMath (external site: huggingface.co) — Math web pages in the Dolma 3 pool.
Catalog records that name this entry in their provenance:
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.