Can data collectives help strengthen vulnerable cultures in the face of AI?
"Data collectives and cooperatives, which let creators control the collection and distribution of their data, are emerging as preferred alternatives to big tech companies."
It’s interesting to contrast the current moment to the “information wants to be free” era of Web 2.0, twenty or so years ago. Back then, everyone was talking about open APIs and open data. Now, it’s become clearer that communities need to control the terms of their data if they’re going to avoid being strip-mined for somebody else’s profit.
“Workers, producers, consumers, and others have been establishing cooperatives and other community-led associations to pool resources, share benefits, and address socioeconomic challenges for centuries. The United Nations marked 2025 as the year of cooperatives, positioning them as “essential solutions to today’s global problems,” kindling renewed interest in data collectives and cooperatives.”
While there’s certainly an argument to be made that communities tend to over-estimate the value of their own data (looking at you, news), some of these datasets may be truly unique in ways that would add value to an AI service or model. As this article points out, collectively-owned data includes creative works in more than 20 African languages that aren’t recognized in mainstream linguistic frameworks.
The danger, of course, is that putting these kinds of gates in front of underrepresented cultures just works to further marginalize them: in that potential future, if everyone’s using a model where those languages are missing, they become irrelevant. But there’s another one where data collectives can pull the levers they have to bring about the world they want to see. That’s exactly what the Nwulite Obodo Open Data License aims to do: data rights holders can negotiate to share their work and cultural heritage without losing their right to benefit from it. (Nwulite Obodo is Igbo for raising, reviving, and building the community.)
In one model, vendors building non-extractive and responsibly trained models for public interest purposes get to use their data for free, but the closed-model big tech vendors have to pay. That’s what Meesum Alam did with voice data for 39 at-risk languages in Pakistan: the communities he worked with determined that the data was free for research and non-commercial purposes, but for-profit tech companies would need to negotiate terms (which Meta did).
That potentially becomes more interesting: either OpenAI et al negotiate to license the data, or they lose functionality to their public interest competitors. There’s also a world where some communities proactively document their cultures and make them available specifically so that models, whoever they’re built by, won’t omit them. Either the world has more equitable AI or the communities financially benefit from their cultural heritage.
Whatever happens, these communities certainly have the right to control their data however they see fit. What vendors do about it is the open question. But initiatives like Mozilla Data Collective make it more possible to have more substantive conversations about how data is provided and used, and that can only be a good thing.