Communities Build Data Collectives to Reclaim Control from Big Tech

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Mozilla Foundation launched Mozilla Data Collective last year, giving communities, organizations, and individuals a platform to publish and govern their own data sets rather than letting Big Tech scrape them freely.
- Meesum Alam uploaded roughly 700 hours of voice recordings spanning 39 endangered Pakistani languages — including Dawoodi (~300 speakers) and Kalasha (~3,000 speakers) — to Mozilla Data Collective, where Meta has since used them to build speech recognition tools.
- Mozilla Data Collective earlier this month released three community-generated data sets for paid commercial licensing, with a broader compensation feature coming so creators are "valued, recognized and supported."
- The Nwulite Obodo Open Data License (NOODL), launched in 2024, now covers about 70 African data sets — including speech data in more than 20 African languages, music, lullabies, and poetry — letting creators share data without surrendering rights to benefit from it.
- Communities on Mozilla Data Collective can restrict firms like Meta to research or non-commercial use and require users to clarify the purpose of access requests, Alam told Rest of World: "They don't trust the big tech companies."
- The United Nations designated 2025 as the International Year of Cooperatives, framing them as "essential solutions to today's global problems" and helping kindle renewed interest in data cooperatives.
Why it matters: Low-resource language communities gain negotiating leverage they never had individually: AI firms like Meta now have to negotiate terms with data collectives rather than scrape freely, and Mozilla's new paid commercial licensing feature formalizes a compensation path for the communities historically excluded from AI development.




