Common Crawl corpus
Version CC-MAIN-2026-39 (latest crawl at review)
Maintained by Common Crawl Foundation12
An archive of web crawl data collected by Common Crawl's crawlers since 2008, released as periodic crawls. Each crawl is published as WARC files (raw HTTP responses, requests, and metadata), WAT files (computed metadata as JSON), and WET files (extracted plain text), together with URL indexes.137
- Documentation: Get started (data access) (external site: commoncrawl.org)
- License: Terms of Use (external site: commoncrawl.org)
- Documentation: FAQ (external site: commoncrawl.org)
- Dataset hub: Crawl index collections (collinfo.json) (external site: index.commoncrawl.org)
Availability and license
Overall availability
Downloadable over HTTPS from data.commoncrawl.org without an AWS account; access to the s3://commoncrawl bucket (AWS us-east-1) through the S3 API is limited to authenticated AWS users. Use is subject to Common Crawl's Terms of Use, and the crawled content may be subject to its original publishers' terms and copyrights.342
Availability is separate from permission: read the license before using or redistributing.
The Terms of Use (last updated March 7, 2024) grant a limited, non-transferable, non-sublicensable license to the service, which is defined to include the crawled content. They state that crawled content is the responsibility of the party it originated from, that it may be subject to separate terms, and that users must respect third parties' copyrights. Users must indemnify Common Crawl against third-party claims arising from their use of the service or crawled content, expressly including use in connection with AI and machine learning systems such as large language models. Disputes go to binding arbitration under California law.2
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| AccessCan the data be obtained, and on what terms? | Public | Downloadable over HTTPS without an AWS account; S3 API access requires an authenticated AWS account.34 |
| ProvenanceAre the data's origins documented? | Public | WARC files keep the HTTP request, response, and crawl metadata for each captured page, and the crawler identifies itself as CCBot and checks robots.txt before fetching.356 |
| DocumentationIs there a datasheet, card, or equivalent documentation? | Public | The Get Started page and FAQ document the formats, indexes, and access methods.36 |
| LicensingAre the licensing terms stated? | Partial | The Terms of Use are published, but no open data license covers the crawled content itself; the Terms say that content may be subject to its owners' separate terms and require users to respect third parties' copyrights.2 |
| Stated limitationsDoes the documentation state known limitations or risks? | Public | The FAQ describes the dataset as a sample of the web that generally archives a randomly selected subset of a site rather than the whole site, and the Terms of Use state that Common Crawl cannot guarantee the accuracy, quality, or lawfulness of the crawled content.62 |
What it is useful for
Organization context
Provenance and derivatives
Catalog records that name this entry in their provenance:
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.