USASI
DatasetDataset

Dolma

Project record

Maintained by Ai2 (Allen Institute for AI)1112

Dolma is Ai2's family of English pretraining corpora built from web pages, academic publications, code, and encyclopedic text, used to train the OLMo models. The first release (v1, August 2023) was updated through v1.7 (April 2024); Dolma 3, used for Olmo 3, consists of a pool of about 9.3 trillion tokens and curated pretraining, mid-training, and long-context mixes. A Dolma 3.5 pool with additional sources and filtering has since been published.129115

Last reviewedEntry updated Documented release Aug 18, 2023

Availability and license

Overall availability

Public

Downloadable from the Hugging Face Hub without gating. The v1.x card states that users are also bound by the license terms of the original data sources, and the Dolma 3 pool repository hosts only its Common Crawl and olmOCR PDF portions, linking the other sources to their original repositories.12

Availability is separate from permission: read the license before using or redistributing.

Dolma moved to ODC-BY in April 2024, per the v1.x dataset card, which also states that users remain bound by the terms of the original data sources. The Dolma 3 cards state the data is intended for research and educational use under Ai2's Responsible Use Guidelines.12

Public materials checklist

Items for a dataset under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for Dolma
ItemStatusNotes and evidence
AccessCan the data be obtained, and on what terms?PublicAll versions are downloadable from Hugging Face without gating; some Dolma 3 pool sources are obtained from their original repositories.12
ProvenanceAre the data's origins documented?PublicDataset cards list sources (for example Common Crawl, code, scientific PDFs, arXiv, math web pages, and Wikipedia) with token and document counts per source.124
DocumentationIs there a datasheet, card, or equivalent documentation?PublicThe Dolma paper includes a datasheet appendix; the dolma3 repository documents how the Dolma 3 datasets were built.619
LicensingAre the licensing terms stated?PublicODC-BY for the data; Apache 2.0 for the toolkit code.128
Stated limitationsDoes the documentation state known limitations or risks?PublicThe paper's Limitations section notes the English-only focus, incomplete coverage of curation practices, ablations run only at 1B scale, and that full manual inspection of the corpus is not feasible. The Olmo 3 7B mix card notes that some PDFs were redacted after training.64

What it is useful for

Documented for language model pretraining and research on pretraining data. The accompanying Dolma toolkit supports tagging, deduplication, mixing, filtering, and tokenization for reproducing or building similar corpora.116

Organization context

Provenance and derivatives

Assembled by Ai2 from existing sources. Dolma v1.7 draws on Common Crawl, RefinedWeb, StarCoder, C4, Reddit, peS2o, arXiv and StackExchange via RedPajama, and Wikipedia, among others. The Dolma 3 pool combines Common Crawl, olmOCR-processed science PDFs, Stack-Edu, arXiv, FineMath, and Wikipedia.12

Catalog records that name this entry in their provenance:

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

Dolma is created and published by Ai2 (Allen Institute for AI) under its Hugging Face and GitHub organizations, and Ai2's documentation presents it as Ai2's training corpus. Ai2 describes itself as a Seattle-based non-profit AI research institute.111712

Assessed Sep 29, 2026

Sources

  1. 1.
    allenai/dolma dataset card (external site: huggingface.co)

    Ai2 · Dataset card · accessed Sep 29, 2026

  2. 2.
    allenai/dolma3_pool dataset card (external site: huggingface.co)

    Ai2 · Dataset card · accessed Sep 29, 2026

  3. 3.
    allenai/dolma3_mix-6T dataset card (external site: huggingface.co)

    Ai2 · Dataset card · accessed Sep 29, 2026

  4. 4.
  5. 5.
    allenai/dolma3.5_pool dataset card (external site: huggingface.co)

    Ai2 · Dataset card · accessed Sep 29, 2026

  6. 6.
    Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research (external site: arxiv.org)

    Ai2 (Soldaini et al.) · Paper · published Jun 6, 2024 · accessed Sep 29, 2026

  7. 7.
    allenai/dolma (external site: github.com)

    Ai2 · Repository · accessed Sep 29, 2026

  8. 8.
  9. 9.
    allenai/dolma3 (external site: github.com)

    Ai2 · Repository · accessed Sep 29, 2026

  10. 10.
  11. 11.
  12. 12.
    About us | Ai2 (external site: allenai.org)

    Ai2 · Official page · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project