The Wikimedia Foundation urges AI giants to stop data plundering and use its paid API

The Wikimedia Foundation urges AI giants to stop data plundering and use its paid API

· updated on 8 August 2026
#wikimedia #techcrunch #theverge

The Wikimedia Foundation has officially requested that artificial intelligence companies stop scraping Wikipedia's data for free, inviting them instead to subscribe to its paid service, Wikimedia Enterprise, in order to support the sustainability of the free encyclopedia and guarantee an ethical use of its content.

The Wikimedia Foundation Demands Compensation from AI Giants for the Use of Wikipedia Data

The Wikimedia Foundation, the non-profit organization behind the online encyclopedia Wikipedia, recently took a firm stance against artificial intelligence (AI) companies. It officially asked them to cease the widespread practice of "scraping" (automated data collection) on Wikipedia and to prioritize access to its content via its paid API, Wikimedia Enterprise. This move marks a turning point in the relationship between the open knowledge ecosystem and the AI industry, raising crucial questions about data value, the sustainability of collaborative projects, and the ethics of commercial exploitation.

Wikipedia: The Goldmine for AI Models

For years, Wikipedia has been an invaluable resource for training large language models (LLMs) and other AI systems. With millions of articles written and updated by volunteers worldwide, the collaborative encyclopedia represents the largest publicly available corpus of structured, multilingual, and verified human knowledge. Its richness, diversity, and constant updates make it an ideal source for teaching AIs the nuances of language, facts, concepts, and the relationships between them.

However, "free and open" access to this content has often been interpreted by AI companies as implicit authorization to scrape massive volumes of data without any direct contribution. This practice, while technically possible, strains Wikimedia's infrastructure and generates no revenue to support the organization's operations, which rely primarily on donations.

The Call for Wikimedia Enterprise: A Mutually Beneficial Solution?

Faced with this situation, the Wikimedia Foundation decided to go on the offensive. It is now urging major AI companies—whose names like Google, OpenAI, Meta, and Microsoft are frequently cited in this context—to adopt a more responsible and sustainable approach. Rather than continuing to "pillage" data, the Foundation invites them to subscribe to Wikimedia Enterprise.

Launched in 2021, Wikimedia Enterprise is a commercial service designed specifically for companies that require large-scale access to Wikipedia and other Wikimedia project data. It provides structured, reliable, real-time data feeds, ensuring higher quality and consistent updates compared to scraped data. The revenue generated by Wikimedia Enterprise is reinvested directly into the Foundation's mission, supporting the infrastructure, technological development, and volunteer communities that make Wikipedia possible.

AI Ethics and the Sustainability of Open Knowledge

This request raises fundamental questions about AI ethics and the sustainability of open resources. As AI companies generate billions of dollars in revenue and market valuation in part by exploiting volunteer-created content, the Wikimedia Foundation believes it is time for these giants to contribute to the sustainability of the source powering their innovations.

Paying for data access via Wikimedia Enterprise is not merely a financial matter; it is also a commitment to data quality and reliability. Companies using this service benefit from an official, maintained source, mitigating risks associated with obsolete or poorly structured data obtained through scraping. Furthermore, it sends a strong message regarding the recognition of the intellectual value and collaborative work that underpins Wikipedia.

What Are the Future Implications?

How AI companies react to this request will be decisive. If they comply, it could set an important precedent for how AI models are trained in the future, encouraging more ethical and sustainable partnerships with content creators and open-source projects. It could also prompt other non-profit platforms to monetize their data to ensure survival against the insatiable appetite of AI.

Conversely, a widespread refusal could escalate tensions, potentially leading to stricter measures from the Wikimedia Foundation—including possible legal actions or tighter technical restrictions on access. The stakes are high: finding a balance between promoting free knowledge and recognizing the economic value it generates for a booming industry.