USASI
DatasetDataset

Common Crawl corpus

Version CC-MAIN-2026-39 (latest crawl at review)

Maintained by Common Crawl Foundation12

An archive of web crawl data collected by Common Crawl's crawlers since 2008, released as periodic crawls. Each crawl is published as WARC files (raw HTTP responses, requests, and metadata), WAT files (computed metadata as JSON), and WET files (extracted plain text), together with URL indexes.137

Last reviewedEntry updated Documented release Sep 2026

Availability and license

Overall availability

Public

Downloadable over HTTPS from data.commoncrawl.org without an AWS account; access to the s3://commoncrawl bucket (AWS us-east-1) through the S3 API is limited to authenticated AWS users. Use is subject to Common Crawl's Terms of Use, and the crawled content may be subject to its original publishers' terms and copyrights.342

Availability is separate from permission: read the license before using or redistributing.

The Terms of Use (last updated March 7, 2024) grant a limited, non-transferable, non-sublicensable license to the service, which is defined to include the crawled content. They state that crawled content is the responsibility of the party it originated from, that it may be subject to separate terms, and that users must respect third parties' copyrights. Users must indemnify Common Crawl against third-party claims arising from their use of the service or crawled content, expressly including use in connection with AI and machine learning systems such as large language models. Disputes go to binding arbitration under California law.2

Public materials checklist

Items for a dataset under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for Common Crawl corpus
ItemStatusNotes and evidence
AccessCan the data be obtained, and on what terms?PublicDownloadable over HTTPS without an AWS account; S3 API access requires an authenticated AWS account.34
ProvenanceAre the data's origins documented?PublicWARC files keep the HTTP request, response, and crawl metadata for each captured page, and the crawler identifies itself as CCBot and checks robots.txt before fetching.356
DocumentationIs there a datasheet, card, or equivalent documentation?PublicThe Get Started page and FAQ document the formats, indexes, and access methods.36
LicensingAre the licensing terms stated?PartialThe Terms of Use are published, but no open data license covers the crawled content itself; the Terms say that content may be subject to its owners' separate terms and require users to respect third parties' copyrights.2
Stated limitationsDoes the documentation state known limitations or risks?PublicThe FAQ describes the dataset as a sample of the web that generally archives a randomly selected subset of a site rather than the whole site, and the Terms of Use state that Common Crawl cannot guarantee the accuracy, quality, or lawfulness of the crawled content.62

What it is useful for

A freely available source of raw and extracted web pages for research and for building web-scale text datasets; for example, Ai2's Dolma pretraining corpus lists Common Crawl as a source.128

Organization context

Provenance and derivatives

Collected from web pages by Common Crawl's own crawlers: a custom Hadoop-based crawler from 2008, replaced in 2013 by the Apache Nutch-based CCBot, which checks robots.txt before fetching. The earliest collection in the index covers 2008 to 2009 (ARC files).1567

Catalog records that name this entry in their provenance:

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The corpus is collected and published by The Common Crawl Foundation, which names itself in the Terms of Use with a Beverly Hills, California address. The IRS lists Commoncrawl Foundation at that address as a 501(c)(3) organization.219

Assessed Sep 29, 2026

Sources

  1. 1.
    About Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Official page · accessed Sep 29, 2026

  2. 2.
    Terms of Use | Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · License · published Mar 7, 2024 · accessed Sep 29, 2026

  3. 3.
    Get Started | Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Documentation · accessed Sep 29, 2026

  4. 4.
  5. 5.
    CCBot | Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Official page · accessed Sep 29, 2026

  6. 6.
    FAQ | Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Documentation · accessed Sep 29, 2026

  7. 7.
    Common Crawl index collections (collinfo.json) (external site: index.commoncrawl.org)

    Common Crawl Foundation · Documentation · accessed Sep 29, 2026

  8. 8.
    allenai/dolma dataset card (external site: huggingface.co)

    Ai2 · Dataset card · accessed Sep 29, 2026

  9. 9.

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project