RedPajama
Version V2
Maintained by Together AI (Together Computer, Inc.)137
RedPajama is Together AI's open pretraining dataset project. RedPajama-V1 (RedPajama-Data-1T, 2023) is a 1.2-trillion-token reproduction of the LLaMA training-data recipe drawn from Common Crawl, C4, GitHub, arXiv, Wikipedia, and Stack Exchange. The current version, RedPajama-V2 (October 2023), contains over 100 billion documents from 84 Common Crawl snapshots in English, German, French, Spanish, and Italian, with more than 40 precomputed quality signals and duplicate markers; its deduplicated, annotated portion is about 30 trillion tokens.21356
- Dataset hub: RedPajama-Data-V2 (Hugging Face) (external site: huggingface.co)
- Dataset hub: RedPajama-Data-1T (Hugging Face) (external site: huggingface.co)
- Repository: RedPajama-Data code repository (external site: github.com)
- Paper: RedPajama paper (arXiv 2411.12372) (external site: arxiv.org)
- Release notes: RedPajama-Data-v2 announcement (external site: together.ai)
Availability and license
Overall availability
Both versions are downloadable from Hugging Face without gating. The V2 card directs users to the Common Crawl Foundation Terms of Use for the data. In V1 the 'book' configuration is marked defunct and no longer accessible because of reported copyright infringement in its Books3 content.12
Availability is separate from permission: read the license before using or redistributing.
Apache License 2.0 (dataset loading and processing code) (external site: raw.githubusercontent.com)41
Neither version grants a single license for the data. The V2 card refers users to the Common Crawl Foundation Terms of Use for the data and applies Apache 2.0 to the code; the V1 card asks users to follow the license of each subset (for example, the GitHub subset was limited to MIT, BSD, and Apache-licensed projects).12
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| AccessCan the data be obtained, and on what terms? | Public | V2 and V1 are downloadable without gating; V1's book subset is no longer accessible.12 |
| ProvenanceAre the data's origins documented? | Public | V2 documents carry their source URL, domain, and Common Crawl snapshot ID, and the card describes the CCNet processing; the V1 card lists each source with its token count and processing steps.12 |
| DocumentationIs there a datasheet, card, or equivalent documentation? | Public | Dataset cards, an announcement post, a NeurIPS 2024 Datasets and Benchmarks paper, and the full pipeline code for recreating V2 (including quality signals) are published.1563 |
| LicensingAre the licensing terms stated? | Partial | Code is Apache 2.0; data terms defer to Common Crawl's terms (V2) or to each subset's license (V1).12 |
| Stated limitationsDoes the documentation state known limitations or risks? | Unknown | Not assessed. |
What it is useful for
Run and use notes
- The V2 card notes that downloading a full snapshot for a given partition and language requires about 1 TB of disk space per snapshot, and provides a small 'sample' configuration for exploration.1
Organization context
Provenance and derivatives
V2 was built from 84 Common Crawl snapshots processed with the CCNet pipeline into head, middle, and tail perplexity buckets; head and middle documents were annotated with quality signals and duplicates were marked with a Bloom filter rather than removed. V1 followed the LLaMA paper's data recipe using Common Crawl dumps processed with CCNet, C4, permissively licensed GitHub code, arXiv, Wikipedia, and Stack Exchange.12
- Derived from: Common Crawl — Sole source for V2 and the largest source for V1.
U.S. eligibility
Eligible · basis: U.S.-governed project
The datasets are published under Together's Hugging Face and GitHub organizations, and the dataset card's software citation names Together Computer as author. Together's terms of service identify Together Computer, Inc., a Delaware corporation, and give a San Francisco, California address for copyright notices. The V2 card also thanks the RedPajama-V1 partners (including Stanford research groups, Mila, Université de Montréal, ETH DS3Lab, LAION, and Ontocord.ai); Together remains the publishing and maintaining entity.137
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.