惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
罗磊的独立博客
宝玉的分享
宝玉的分享
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
V
V2EX
酷 壳 – CoolShell
酷 壳 – CoolShell
T
Tailwind CSS Blog
博客园_首页
量子位
月光博客
月光博客
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 司徒正美
人人都是产品经理
人人都是产品经理
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
爱范儿
爱范儿
S
SegmentFault 最新的问题
雷峰网
雷峰网
小众软件
小众软件
博客园 - 聂微东
美团技术团队
Apple Machine Learning Research
Apple Machine Learning Research
WordPress大学
WordPress大学
Jina AI
Jina AI
Hugging Face - Blog
Hugging Face - Blog

Rest of World -

Why we should all be worried about AI in elections How one VC burns through hundreds of millions of tokens a day to find the next unicorn Beijing is forcing a mass breakup with AI lovers Arizona wants Taiwan’s investors to think beyond chips The offline messaging apps challenging internet shutdowns Growth without work: The human cost of the AI revolution The global grassroots gatherings trying to humanize the AI boom Why Silicon Valley is divided over China’s powerful, cheap AI models Can we train AI to choose safety over speed? With Moonshot’s free Kimi K3, China changes the sovereign AI playbook Anxious Chinese students are trusting AI to help pick colleges and majors Indian EV makers beat Tesla and BYD on energy efficiency In China, people are renting out their faces to AI Why state-owned AI won’t solve inequality The hidden cost of Myanmar’s rare earth mines The U.S. wants to contain China’s AI. Silicon Valley keeps using it China’s AI talent race is starting in high school AI is shrinking video game development teams to one Can AI beat a goldfish at calling the World Cup? The problem AI content moderation cannot solve AI powers citizen-led disaster relief from afar for Venezuela The Gulf has billions to spend on AI. It still needs Nvidia India’s crackdown on a new WhatsApp feature risks setting a global precedent Older adults know AI is slop. They just like it Your next nurse may monitor you from the Philippines Data centers should benefit the cities that power them China’s AI boom is creating a different kind of entrepreneur China’s web novel platforms embraced AI. Now they are fighting it India is testing an alternative to Silicon Valley’s AI playbook China’s EV makers are taking over the European factories Ford and Nissan can’t fill
Fed up with Big Tech, communities turn to data collective...
Christine Glancey · 2026-07-23 · via Rest of World -

A growing pushback against big tech companies, and greater awareness of the value of data, is spurring interest in data collectives and cooperatives, which give communities control over the collection, management, and distribution of their data. This alternative allows creators to benefit from data sets that may otherwise be ignored or misused.

A handful of tech companies dominate the generative artificial intelligence industry, with American and Chinese frontier models controlling the lion’s share of the market. Some countries are building their own large language models because their language and culture are not adequately represented in GPT, Gemini, Claude, or Qwen.

The big companies have built themselves up on the backs of all these people creating data, who think it’s time to set their own terms now.”

Raffi Krikorian, chief technology officer, Mozilla Foundation

Communities that possess smaller or unusual data sets can gain from having control over them, Raffi Krikorian, chief technology officer at Mozilla Foundation, told Rest of World. The nonprofit last year set up Mozilla Data Collective to provide a platform for such data sets from communities, organizations, and individuals around the world.

“More people are starting to feel like they’re sitting on unique information, not generally available on the internet, and they want to turn the tables on what governance looks like for that data,” Krikorian said.

“The anti-Big AI, anti-Big Tech push is a convenient bedfellow. The big companies have built themselves up on the backs of all these people creating data, who think it’s time to set their own terms now,” he said.

A blueprint for responsible use

Companies including Meta, Open AI, Google, and Anthropic have scraped nearly all available data from the internet to train their AI models, and have been accused of using copyrighted material without consent. Tech companies have said it qualifies as fair use, which allows the use of such material for research and other purposes. Some countries are trying to balance the need for good data with the need to compensate creators.

Data collectives, or cooperatives, offer a blueprint for the responsible use of data, and ensure the value goes to those generating the data, Astha Kapoor, co-founder and director of Aapti Institute, a tech research firm in India, told Rest of World.

“Beyond consent and compensation, collective action around data gives communities the opportunity to direct data towards issues they may care about,” she said. “Communities can negotiate the terms on which their data is used at every stage of the AI lifecycle [with] mechanisms for accountability and redressal, in case their terms are breached.”

Workers, producers, consumers, and others have been establishing cooperatives and other community-led associations to pool resources, share benefits, and address socioeconomic challenges for centuries. The United Nations marked 2025 as the year of cooperatives, positioning them as “essential solutions to today’s global problems,” kindling renewed interest in data collectives and cooperatives.

They cover a wide range: The Kerala Food Platform enables about 2,500 farmers in the southern Indian state to trace and market their produce, including rice, fish, fruits, and vegetables. Mexico-based PescaData helps small-scale fishers in Latin America and the Caribbean to manage and benefit from their catch records, while the Native BioData Consortium is a repository of the genetic and environmental data of Indigenous people.

Increasingly, data sets are being created for AI-related purposes, including in low-resource language communities, whose data, including voice data, is valuable for training small models and creating speech recognition tools. For these communities, a data collective is a more practical solution, Krikorian said.

“An OpenAI or Anthropic is not going to prioritize a language that’s only spoken by 1 million people,” he said. “But if it existed, a chatbot that can communicate in their language is hugely beneficial to that community.” 

“Linguistic identity crisis”

Long before the launch of ChatGPT, Meesum Alam realized that dozens of languages were dying in his native Pakistan. He belonged to the Baloch community but could not speak Balochi, the language of his forefathers, and experienced a “linguistic identity crisis,” he told Rest of World.

Every other day, I get a text or a voice note from someone who is able to communicate with a chatbot in their own language for the first time.”Meesum Alam, Ph.D. candidate in computational linguistics at Indiana University

Alam began to document languages that were at risk of dying because few people spoke them. He began with Dawoodi, which had about 300 speakers, and Kalasha, with some 3,000 speakers, and collected voice data in 39 languages from communities and local organizations in Pakistan. He put the data sets, totaling about 700 hours, on Mozilla Data Collective, where they have been used by companies including Meta to build speech recognition tools, he said.

“These languages were never part of the digital world,” said Alam, a Ph.D. candidate in computational linguistics at Indiana University. “For the communities, being able to interact with AI in their own language is a first step into the AI world.” 

The communities decided their data sets can only be used for research or non-commercial purposes. Meta and other big companies have to negotiate the terms of use with the communities, Alam said. Money is not a motivating factor, he said. 

“They don’t trust the big tech companies because they know they can take the data and monetize it,” he said. “They want a fair deal for the entire community, and they can only get that through a data collective. It gives some power back to communities.”

Besides collectives, other frameworks in use include data trusts, where a trustee manages data on behalf of a group; data unions, which aggregate the data of individual members to negotiate collectively with buyers; data commons, such as Wikimedia and OpenStreetMap, which have different governance structures; and data donation schemes, such as the Personal Genome Project, where individuals contribute their data for public benefit.

“It’s a very emotional thing”

In the African continent, which has long been subject to extractive practices, the Nwulite Obodo Open Data License, launched in 2024, enables creators, communities, researchers, and others to share data sets without giving up the right to benefit from them. About 70 African data sets under NOODL are part of the Mozilla Data Collective, including speech data sets in more than 20 African languages, music, lullabies, and poetry.

These languages are not recognized in the mainstream linguistic framework, so having them in the data collective increases their visibility and accessibility, Emmanuel Ngue Um, a regional researcher at the Institute of African Digital Humanities, which has published about 40 data sets on Mozilla Data Collective, told Rest of World.

“Most importantly, they can require users to clarify the purpose of their access request,” he said. “This creates opportunities for cooperation and partnership between those with the technological ability to support language work in underserved communities, and the communities.”

Still, data collectives can face governance and scaling challenges, Kapoor said. 

“Building data cooperatives solely to steward data is not feasible because sustainability becomes an issue, and the only viable pathway becomes monetization of the data the cooperative is meant to safeguard, which is problematic,” she said.

Earlier this month, the Mozilla Data Collective made three community-generated data sets available for paid commercial licensing, ahead of opening the compensation feature to all users so that communities creating the data are “valued, recognized and supported,” it said.

For Alam, who is on the hunt for low-resource languages in India and Bangladesh as well, the benefit is clear. AI adoption is growing quickly, and for many communities, having their language data sets easily accessible through a data collective is the only way they can participate.  

“Every other day, I get a text or a voice note from someone who is able to communicate with a chatbot in their own language for the first time,” he said. “It’s a very emotional thing for them when they can do that.”