USASI

Organization

Common Crawl

Data services · Beverly Hills, California (address given in the Terms of Use and the IRS listing)

Common Crawl is a nonprofit that crawls the web and makes its archives and derived datasets freely available. It has collected crawl data since 2008 and hosts it on Amazon Web Services through an AWS open data sponsorship program.13

Last reviewedEntry updated

Key facts

Legal name
The Common Crawl Foundation (listed by the IRS as Commoncrawl Foundation)26
Ownership
Nonprofit
Legal form
Corporation recognized by the IRS as tax-exempt under Section 501(c)(3); the IRS extract classifies it as a private operating foundation.67
Headquarters
Beverly Hills, California (address given in the Terms of Use and the IRS listing)26
Founded
20071
Sectors
Data

What is and isn’t public

An evidence-based summary of what this catalog has documented. It is not a verdict on the organization as a whole.

Publishes its crawl archives (WARC, WAT, and WET files) and indexes for download over HTTPS without an AWS account. Use is governed by Common Crawl's Terms of Use rather than a named open data license, and the crawled content remains subject to the rights of its original publishers.32

Products

Documented products and how they are delivered. Each row links to the official page.

No product records yet.

Related open artifacts

Records in the Open Models & Tools directory that list this organization as a maintainer or publisher.

Datasets (1)

Notable documented facts

  • Common Crawl's crawler, CCBot, checks robots.txt before fetching pages and honors the Crawl-delay parameter and nofollow attributes; site owners can block it through robots.txt, and Common Crawl also refers site owners to an opt-out registry.45
  • The Terms of Use (last updated March 7, 2024) require users to indemnify Common Crawl against third-party claims arising from their use of the service or crawled content, expressly including use in connection with artificial intelligence and machine learning systems such as large language models.2

U.S. eligibility

How this record meets the catalog’s published eligibility policy.

Eligible · basis: U.S. nonprofit or lab

Common Crawl describes itself as a 501(c)(3) nonprofit. Its Terms of Use name The Common Crawl Foundation with a Beverly Hills, California address, and the IRS exempt-organization extract for California lists Commoncrawl Foundation at the same address as a 501(c)(3) organization.126

Assessed Sep 29, 2026

Eligibility policy

Sources

  1. 1.
    About Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Official page · accessed Sep 29, 2026

  2. 2.
    Terms of Use | Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Official page · published Mar 7, 2024 · accessed Sep 29, 2026

  3. 3.
    Get Started | Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Documentation · accessed Sep 29, 2026

  4. 4.
    CCBot | Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Official page · accessed Sep 29, 2026

  5. 5.
    FAQ | Common Crawl (external site: commoncrawl.org)

    Common Crawl Foundation · Documentation · accessed Sep 29, 2026

  6. 6.
  7. 7.
    Exempt Organizations Business Master File Extract (EO BMF) (external site: irs.gov)

    Internal Revenue Service · Documentation · published Jan 2026 · accessed Sep 29, 2026

This listing is not an endorsement or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project