惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
Stack Overflow Blog
Stack Overflow Blog
B
Blog RSS Feed
C
Check Point Blog
D
Docker
Y
Y Combinator Blog
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
MongoDB | Blog
MongoDB | Blog
博客园_首页
Apple Machine Learning Research
Apple Machine Learning Research
量子位
有赞技术团队
有赞技术团队
IT之家
IT之家
大猫的无限游戏
大猫的无限游戏
D
DataBreaches.Net
M
MIT News - Artificial intelligence
B
Blog
阮一峰的网络日志
阮一峰的网络日志
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
腾讯CDC
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
V
V2EX
月光博客
月光博客

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, and November 2025
Common Crawl - Blog - A Sampling of 2025 Research Referen...
2025-12-05 · via Common Crawl

As another year here at Common Crawl comes to a close, we present a dozen papers (selected from the thousands published in 2025) that demonstrate the range of topics and areas of study for which Common Crawl’s datasets and statistics are used and referenced. For more papers citing Common Crawl’s data, which has been regularly collected since 2008, see Research Papers and cc-citations, our curated BibTeX database.

Classification of Worldwide News Articles by Perceived Quality, 2018-2024

“This study explored whether supervised machine learning and deep learning models can effectively distinguish perceived lower-quality news articles from perceived higher-quality news articles. 3 machine learning classifiers and 3 deep learning models were assessed using a newly created dataset of 1,412,272 English news articles from the Common Crawl over 2018-2024.”

https://arxiv.org/abs/2511.16416

Combating Health Misinformation With Fusion-Based Credible Retrieval Techniques

“This study aims to combat health misinformation by enhancing the retrieval of credible health information using effective fusion-based techniques. … The datasets for these events are based on the CommonCrawl News dataset”

https://journals.sagepub.com/doi/10.1177/14604582251388860

Geospatiality: The Effect of Topics on the Presence of Geolocation in English Text Data

“This study investigates the relationship between texts’ thematic categories and their likelihood of containing usable geolocation information by quantifying and modelling this relationship across seven diverse English text datasets of different types, including web forums, microblogs, news, and magazines.” Study uses Common Crawls Distribution of Languages statistics.

https://www.tandfonline.com/doi/full/10.1080/13658816.2025.2460051#abstract

High-Fidelity Simultaneous Speech-To-Speech Translation

Introduces  Hibiki, “a decoder-only model for simultaneous speech translation.”  The “training dataset is made of filtered web pages from Common Crawl, as well as curated sources such as Wikipedia, StackExchange or scientific articles and it contains 12.5% of multilingual documents.”

https://arxiv.org/abs/2502.03382

Optimising Web Accessibility Evaluation: Population Sourcing Methods for Web Accessibility Evaluation

“We present a tool-supported framework, OPTIMAL-EM, that runs parallel to the Website Accessibility Conformance Evaluation Methodology (WCAG-EM). We aim to optimise web accessibility evaluation through the targeted use of automated tools and human evaluation to audit a more representative set of pages.”  Uses four approaches, including Common Crawl.

https://www.sciencedirect.com/science/article/pii/S1071581925000291?via%3Dihub

Paraphrase Detection for Urdu Language Text Using Fine-Tune BiLSTM Framework

“This research proposes a novel bidirectional long short-term memory (BiLSTM) framework to address Urdu paraphrase detection’s intricacies.“ … The study “incorporates the GloVe approach for embedding words with 50 dimensions. [It uses] Common Crawl pre-trained vectors trained on a large amount of web-based text (42 billion tokens, 1.9 million words, 50 d vectors). “

https://www.nature.com/articles/s41598-025-93260-6

Reinforced Disentangled HTML Representation Learning with Hard-Sample Mining for Phishing Webpage Detection

“This study introduces a reinforced Triplet Network to optimize disentangled representation learning tailored for phishing detection…The datasets used in this study include benign data from Common Crawl and phishing data from Phishtank and Mendeley Data.”

https://www.mdpi.com/2079-9292/14/6/1080

Semantic Annotation Model and Method Based on Internet Open Dataset

“[T]his paper deeply studies the semantic annotation model and method based on internet open datasets, aiming to improve annotation efficiency and accuracy and promote data resource sharing and utilization. This paper selects Common Crawl dataset to provide sufficient training samples; methods such as removing stop words and deduplication are used to preprocess data to improve data quality; a keyword extraction model based on heuristic rules and text context is constructed.”

https://www.igi-global.com/gateway/article/370966

Scalable Private Partition Selection via Adaptive Weighting

Proposes “an algorithm for this problem, MaxAdaptiveDegree (MAD), which adaptively reroutes weight from items with weight far above the threshold needed for privacy to items with smaller weight, thereby increasing the probability that less frequent items are output.” [Uses Common Crawl datasets]

https://arxiv.org/abs/2502.08878

SocialQuotes: Learning Contextual Roles of Social Media Quotes on the Web

Introduces SocialQuotes, “a new data set built from the Common Crawl of over 32 million social quotes, 8.3k of them with crowdsourced quote annotations.”

https://ojs.aaai.org/index.php/ICWSM/article/view/35882

Temporally Extending Existing Web Archive Collections for Longitudinal Analysis

This paper introduces “a methodology to extend existing web archive collections temporally to enable longitudinal analysis, including a dataset extended with this methodology [to identify] reasons URL candidates could be missing from the … [Environmental Governance and Data Initiative] EDGI dataset, and crawled the past web of 2008 in order to identify these missing pages.” Includes Common Crawl, in addition to Internet Archive and the End of Term Archive.

https://arxiv.org/abs/2505.24091

Web2Wiki: Characterizing Wikipedia Linking Across the Web

Presents “the first large-scale analysis of how Wikipedia is referenced across the Web. Using a dataset from Common Crawl [it identifies] over 90 million Wikipedia links spanning 1.68% of Web domains and examine their distribution, context, and function.”

https://arxiv.org/abs/2505.15837

With over 10,000 research papers referencing Common Crawl’s dataset, we are constantly surprised by the myriad ways in which our dataset is useful for academic research. We look forward to being surprised again in 2026!