惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Troy Hunt's Blog
P
Palo Alto Networks Blog
N
News and Events Feed by Topic
Hacker News - Newest:
Hacker News - Newest: "LLM"
T
Threatpost
C
Cyber Attacks, Cyber Crime and Cyber Security
S
Schneier on Security
Google Online Security Blog
Google Online Security Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Spread Privacy
Spread Privacy
NISL@THU
NISL@THU
Cisco Talos Blog
Cisco Talos Blog
The GitHub Blog
The GitHub Blog
S
SegmentFault 最新的问题
量子位
L
Lohrmann on Cybersecurity
酷 壳 – CoolShell
酷 壳 – CoolShell
Attack and Defense Labs
Attack and Defense Labs
Y
Y Combinator Blog
Project Zero
Project Zero
AWS News Blog
AWS News Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Last Week in AI
Last Week in AI
博客园 - 聂微东
MyScale Blog
MyScale Blog
aimingoo的专栏
aimingoo的专栏
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
S
Securelist
Latest news
Latest news
C
CXSECURITY Database RSS Feed - CXSecurity.com
B
Blog RSS Feed
Webroot Blog
Webroot Blog
Blog — PlanetScale
Blog — PlanetScale
Recent Announcements
Recent Announcements
V2EX - 技术
V2EX - 技术
Schneier on Security
Schneier on Security
F
Full Disclosure
Apple Machine Learning Research
Apple Machine Learning Research
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
P
Proofpoint News Feed
Recent Commits to openclaw:main
Recent Commits to openclaw:main
月光博客
月光博客
L
LINUX DO - 最新话题
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
H
Heimdal Security Blog
F
Fortinet All Blogs
博客园_首页
N
News | PayPal Newsroom
P
Proofpoint News Feed

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, and November 2025 Common Crawl - Blog - November 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Celebrates World Digital Preservation Day Common Crawl - Blog - Setting the Record Straight: Common Crawl’s Commitment to Transparency, Fair Use, and the Public Good Common Crawl - Blog - October/November 2025 Newsletter Common Crawl - Blog - Common Crawl Foundation at Stanford HAI Common Crawl - Blog - Host- and Domain-Level Web Graphs August, September, and October 2025 Common Crawl - Blog - October 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at COLM 2025 Common Crawl - Blog - Announcing GneissWeb Annotations Common Crawl - Blog - Web Languages Needing Review by Native Speakers Common Crawl - Blog - Host- and Domain-Level Web Graphs July, August, and September 2025 Common Crawl - Blog - From SEO to AIO: Why Your Content Needs to Exist in AI Training Data Common Crawl - Blog - September 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation Opt-Out Registry Common Crawl - Blog - Trip Report: AI_dev (Linux Foundation) August 2025 Common Crawl - Blog - Common Crawl Foundation at Stanford HAI: A Shared Legacy of Data and Innovation Common Crawl - Blog - July/August 2025 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs June, July, and August 2025 Common Crawl - Blog - August 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at ACL 2025 Common Crawl - Blog - AI Optimization Is Here: Are You Ready for Search 2.0? Common Crawl - Blog - IETF 123 Report Common Crawl - Blog - Host- and Domain-Level Web Graphs May, June, and July 2025 Common Crawl - Blog - July 2025 Crawl Archive Now Available Common Crawl - Blog - WMDQS Shared Task on Language Identification Common Crawl - Blog - The First WMDQS-Masakhane LangID Hackathon Common Crawl - Blog - Host- and Domain-Level Web Graphs April, May, and June 2025 Common Crawl - Blog - Common Crawl at the United Nations Open Source Week, June 2025 Common Crawl - Blog - June 2025 Crawl Archive Now Available Common Crawl - Blog - May/June 2025 Newsletter Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets using Python Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2025 Common Crawl - Blog - May 2025 Crawl Archive Now Available Common Crawl - Blog - Announcing the First Workshop on Multilingual Data Quality Signals Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2025 Common Crawl - Blog - April 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing the Host Index Common Crawl - Blog - IIPC General Assembly & Web Archiving Conference 2025 Common Crawl - Blog - March/April 2025 Newsletter Common Crawl - Blog - Providing Authenticity & Data Provenance for Common Crawl Using Blockchain: Our Work with Constellation Network Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2025 Common Crawl - Blog - March 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing Common Crawl AI Agent by ReadyAI Common Crawl - Blog - Submission to the UK’s Copyright and AI Consultation Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2024 and January/February 2025 Common Crawl - Blog - February 2025 Crawl Archive Now Available Common Crawl - Blog - Opening the Gates to Online Safety Common Crawl - Blog - January/February 2025 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2024 and January 2025 Common Crawl - Blog - January 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing cc-downloader Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, and December 2024 Common Crawl - Blog - Common Crawl Foundation at NeurIPS 2024: Expanding Horizons and Building Connections Common Crawl - Blog - Expanding the Language and Cultural Coverage of Common Crawl Common Crawl - Blog - October/November 2024 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, November 2024 Common Crawl - Blog - November 2024 Crawl Archive Now Available Common Crawl - Blog - Reflections on Recent Talks at the Turing Institute and UCL Common Crawl - Blog - Introducing the Common Crawl Errata Page for Data Transparency Common Crawl - Blog - Host- and Domain-Level Web Graphs August, September, and October 2024 Common Crawl - Blog - October 2024 Crawl Archive Now Available Common Crawl - Blog - White House Briefing on Open Data’s Role in Technology Common Crawl - Blog - IAB Workshop on AI-CONTROL Common Crawl - Blog - Host- and Domain-Level Web Graphs July, August, and September 2024 Common Crawl - Blog - September 2024 Crawl Archive Now Available Common Crawl - Blog - August/September 2024 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs June, July, and August 2024 Common Crawl - Blog - August 2024 Crawl Archive Now Available Common Crawl - Blog - The Increase of Common Crawl Citations in Academic Research
Common Crawl - Blog - December 2024 Crawl Archive Now Available
2024-12-18 · via Common Crawl

The crawl archive for December 2024 is now available.

The data was crawled between December 1st and December 15th, and contains 2.64 billion web pages (or 394 TiB of uncompressed content). Page captures are from 47.5 million hosts or 38.3 million registered domains and include 1.05 billion new URLs, not visited in any of our prior crawls.

Archive Location & Download

The December 2024 crawl archive is located in the commoncrawl bucket with the prefix: crawl-data/CC-MAIN-2024-51/.

To assist with exploring and using the dataset, we provide gzip-compressed files which list all segments, WARC, WAT and WET files.

By simply adding either s3://commoncrawl/ or https://data.commoncrawl.org/ to each line, you end up with the S3 and HTTP paths respectively. Please see Get Started for detailed instructions.

Changes to the WAT Metadata Format

Multi-valued headers

Repeated HTTP and WARC headers were not represented in the JSON data in WAT files. When a header was repeated adding a further value of that header, only the last value was stored and other values were lost. This old issue (ia-web-commons#18) is now fixed:

  • Single value headers are represented as before by a header name and a string value.
  • Headers with multiple values are represented by a header name and an associated list of values.

Users are advised to update any code consuming WAT files to this change. The examples in the projects cc-pyspark and cc-warc-examples were updated accordingly, see cc-pyspark#46 resp. cc-warc-examples#5.

Below are two JSON snippets of multi-valued headers:

{
  "Container": { "...": "..." },
  "Envelope": {
        "WARC-Header-Metadata": {
          "...": "...",
          "WARC-Target-URI": "https://en.wikipedia.org/wiki/Saturn",
          "WARC-Protocol": [
            "h2",
            "tls/1.3"
          ],
  • Many HTTP headers, most commonly the "Set-Cookie" header:
{
"Container": { "...": "..." },
"Envelope": {
  "Payload-Metadata": {
    "Actual-Content-Type": "application/http; msgtype=response",
    "HTTP-Response-Metadata": {
      "...": "...",
      "Headers": {
        "date": "Sat, 30 Nov 2024 11:13:30 GMT",
        "...": "...",
        "set-cookie": [
          "WMF-Last-Access=30-Nov-2024;Path=/;HttpOnly;secure;Expires=Wed, 01 Jan 2025 12:00:00 GMT",
          "WMF-Last-Access-Global=30-Nov-2024;Path=/;Domain=.wikipedia.org;HttpOnly;secure;Expires=Wed, 01 Jan 2025 12:00:00 GMT",
          "WMF-DP=5b0;Path=/;HttpOnly;secure;Expires=Sun, 01 Dec 2024 00:00:00 GMT",
          "GeoIP=US:VA:Ashburn:39.05:-77.49:v4; Path=/; secure; Domain=.wikipedia.org",
          "NetworkProbeLimit=0.001;Path=/;Secure;SameSite=Lax;Max-Age=3600"
        ],

Add language attributes of the <html> root element as metadata

The WAT metadata now includes the language attributes of the <html> element. For example, the root element <html lang="es-MX"> is stored in the WAT file as:

"HTML-Metadata": {
"Head": {
    "Metas": [
         {
        "name": "HTML@/lang",
        "content": "en"
         },

Details on this change are tracked in ia-web-commons#35.

Do not include <meta itemprop="..."> as metadata

Schema.org annotations in <meta itemprop="..."> in the HTML body are not put as metadata into the WAT metadata, cf. ia-web-commons#40.

Crawling with IPv6

The crawler is now ready to crawl IPv6-only websites. While IPv4 is still preferred, sites which are only available by IPv6 are now visited by our crawler. As a consequence, IPv6 addresses now appear in the crawl data. For example, in the "WARC-IP-Address" header or in URLs in the URL indexes.

Crawler Verification

Our crawler "CCBot" is now run on dedicated IP address ranges with reverse DNS. This allows webmasters to verify whether a logged request stems from CCBot. Please read our FAQ for more information.

Feedback Welcome

We look forward to hearing your thoughts and comments. As ever, please feel free to join the discussions in our Google Group or in our Discord server.