USASI
DatasetDataset

RedPajama

Version V2

Maintained by Together AI (Together Computer, Inc.)137

RedPajama is Together AI's open pretraining dataset project. RedPajama-V1 (RedPajama-Data-1T, 2023) is a 1.2-trillion-token reproduction of the LLaMA training-data recipe drawn from Common Crawl, C4, GitHub, arXiv, Wikipedia, and Stack Exchange. The current version, RedPajama-V2 (October 2023), contains over 100 billion documents from 84 Common Crawl snapshots in English, German, French, Spanish, and Italian, with more than 40 precomputed quality signals and duplicate markers; its deduplicated, annotated portion is about 30 trillion tokens.21356

Last reviewedEntry updated Documented release Oct 30, 2023

Availability and license

Overall availability

Public

Both versions are downloadable from Hugging Face without gating. The V2 card directs users to the Common Crawl Foundation Terms of Use for the data. In V1 the 'book' configuration is marked defunct and no longer accessible because of reported copyright infringement in its Books3 content.12

Availability is separate from permission: read the license before using or redistributing.

Neither version grants a single license for the data. The V2 card refers users to the Common Crawl Foundation Terms of Use for the data and applies Apache 2.0 to the code; the V1 card asks users to follow the license of each subset (for example, the GitHub subset was limited to MIT, BSD, and Apache-licensed projects).12

Public materials checklist

Items for a dataset under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for RedPajama
ItemStatusNotes and evidence
AccessCan the data be obtained, and on what terms?PublicV2 and V1 are downloadable without gating; V1's book subset is no longer accessible.12
ProvenanceAre the data's origins documented?PublicV2 documents carry their source URL, domain, and Common Crawl snapshot ID, and the card describes the CCNet processing; the V1 card lists each source with its token count and processing steps.12
DocumentationIs there a datasheet, card, or equivalent documentation?PublicDataset cards, an announcement post, a NeurIPS 2024 Datasets and Benchmarks paper, and the full pipeline code for recreating V2 (including quality signals) are published.1563
LicensingAre the licensing terms stated?PartialCode is Apache 2.0; data terms defer to Common Crawl's terms (V2) or to each subset's license (V1).12
Stated limitationsDoes the documentation state known limitations or risks?UnknownNot assessed.

What it is useful for

Building filtered pretraining mixes for large language models: users select documents by snapshot, language, and partition and apply the provided quality signals and duplicate IDs to create their own subsets.16

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The V2 card notes that downloading a full snapshot for a given partition and language requires about 1 TB of disk space per snapshot, and provides a small 'sample' configuration for exploration.1

Organization context

Provenance and derivatives

V2 was built from 84 Common Crawl snapshots processed with the CCNet pipeline into head, middle, and tail perplexity buckets; head and middle documents were annotated with quality signals and duplicates were marked with a Bloom filter rather than removed. V1 followed the LLaMA paper's data recipe using Common Crawl dumps processed with CCNet, C4, permissively licensed GitHub code, arXiv, Wikipedia, and Stack Exchange.12

  • Derived from: Common Crawl — Sole source for V2 and the largest source for V1.

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The datasets are published under Together's Hugging Face and GitHub organizations, and the dataset card's software citation names Together Computer as author. Together's terms of service identify Together Computer, Inc., a Delaware corporation, and give a San Francisco, California address for copyright notices. The V2 card also thanks the RedPajama-V1 partners (including Stanford research groups, Mila, Université de Montréal, ETH DS3Lab, LAION, and Ontocord.ai); Together remains the publishing and maintaining entity.137

Assessed Sep 29, 2026

Sources

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project