Organization
Common Crawl
Data services · Beverly Hills, California (address given in the Terms of Use and the IRS listing)
Common Crawl is a nonprofit that crawls the web and makes its archives and derived datasets freely available. It has collected crawl data since 2008 and hosts it on Amazon Web Services through an AWS open data sponsorship program.13
Key facts
What is and isn’t public
Products
No product records yet.
Related open artifacts
Datasets
- Common Crawl corpusDatasetAvailability: Public
Notable documented facts
- Common Crawl's crawler, CCBot, checks robots.txt before fetching pages and honors the Crawl-delay parameter and nofollow attributes; site owners can block it through robots.txt, and Common Crawl also refers site owners to an opt-out registry.45
- The Terms of Use (last updated March 7, 2024) require users to indemnify Common Crawl against third-party claims arising from their use of the service or crawled content, expressly including use in connection with artificial intelligence and machine learning systems such as large language models.2
U.S. eligibility
Eligible · basis: U.S. nonprofit or lab
Common Crawl describes itself as a 501(c)(3) nonprofit. Its Terms of Use name The Common Crawl Foundation with a Beverly Hills, California address, and the IRS exempt-organization extract for California lists Commoncrawl Foundation at the same address as a 501(c)(3) organization.126
Sources
This listing is not an endorsement or a federal approval.