惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

量子位
F
Fortinet All Blogs
J
Java Code Geeks
Y
Y Combinator Blog
Stack Overflow Blog
Stack Overflow Blog
V
Visual Studio Blog
M
MIT News - Artificial intelligence
腾讯CDC
Last Week in AI
Last Week in AI
The Cloudflare Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
Jina AI
Jina AI
Microsoft Security Blog
Microsoft Security Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
P
Proofpoint News Feed
博客园 - 叶小钗
Recent Announcements
Recent Announcements
T
Tailwind CSS Blog
Engineering at Meta
Engineering at Meta
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
人人都是产品经理
人人都是产品经理
L
LangChain Blog
博客园 - 司徒正美
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - Opening the Gates to Online Safety
2025-02-17 · via Common Crawl

ROOST: Robust Online Open Safety Tools

Last week in Paris, at the AI Action Summit, a coalition of major technology companies and foundations announced the launch of ROOST: Robust Online Open Safety Tools (https://roost.tools). ROOST makes critical data and tools for online safety openly accessible to benefit everyone; a mission which closely aligns with ours at Common Crawl.

The lack of robust infrastructure for online safety [1] has had significant consequences. For example, research [2] [3] has shown that large language models (LLMs) generate significantly more unsafe responses in non-English languages than in English, a disparity which Common Crawl's recent efforts to improve coverage of low-resource languages aim to address, but initiatives like ROOST further bridge the gap in infrastructure by providing accessible safety tools for a wider range of contexts.

Work to improve AI safety across the industry has only just begun.

“Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of potential existential threats. However, this framing has three potential drawbacks: it may exclude researchers and practitioners who are committed to AI safety but approach the field from different angles; it could lead the public to mistakenly view AI safety as focused solely on existential scenarios rather than addressing a wide spectrum of safety challenges; and it risks creating resistance to safety measures among those who disagree with predictions of existential AI risks.”
~ AI Safety for Everyone, Balint Gyevnar, et al, February 2025 [4]

AI safety is often discussed in broad theoretical terms, but practical solutions (tools, resources, and frameworks) are often closed off, expensive, or controlled by a few major players. This not only reduces effectiveness of safety interventions, but also creates barriers for smaller organisations and independent developers. ROOST aims to ensure that developers at all levels can implement best practices in safety.

Left to right: Juliet Shen, Emily Liu, Clint Smith, Audrey Tang, Thom Vaughan, Vilas Dhar, Camille François, Eli Sugarman, Chris DiBona, Paul Ash, Alexandra Reeve Givens, Nabiha Syed, at the ROOST launch in Paris, France.

We at Common Crawl have always believed that access to high quality web data should not be limited to a select few. The Internet is a shared resource, and making web data freely available has driven incredible progress in innovation and across countless research fields. In the same way, opening up safety tools creates a much healthier and more balanced ecosystem in tech, where developers, researchers, and policymakers can work together to build safer and more transparent systems.

References

[1] "Landscape of AI safety concerns -- A methodology to support safety assurance for AI-based autonomous systems", Ronald Schnitzer, et al. https://arxiv.org/abs/2412.14020

[2] "All Languages Matter: On the Multilingual Safety of Large Language Models", Wenxuan Wang, et al. https://arxiv.org/abs/2310.00905

[3] "Multilingual Jailbreak Challenges in Large Language Models", Yue Deng, et al. https://arxiv.org/html/2310.06474v3

[4] "AI Safety for Everyone", Balint Gyevnar, et al. https://arxiv.org/abs/2502.09288