Technology
Ars Technica - All content

Microsoft exec called AI scraping the “largest theft of labor in human history”

Source Entity

Ashley Belanger

September 20, 2026
Microsoft exec called AI scraping the “largest theft of labor in human history”

Unsealed court filings reveal Microsoft and OpenAI executives privately acknowledged that their AI data scraping practices amounted to 'theft' and posed an existential threat to journalism. These documents expose deliberate efforts to bypass paywalls and strip copyright notices while training AI models.

The Hidden Conflict: AI Training and Intellectual Property

Newly unsealed court documents from the ongoing copyright litigation between The New York Times and the tech giants OpenAI and Microsoft have pulled back the curtain on the internal anxieties surrounding AI development. The filings expose a striking contradiction: while publicly championing AI as a tool for progress, internal communications reveal that leadership at both companies understood their data acquisition methods were not only controversial but potentially destructive to the very foundations of the journalism industry.

The Admission of 'Theft'

At the heart of the disclosure is the candid admission by Microsoft Director of Applied Science Brent Hecht, who characterized the mass scraping of news content as the “largest theft of labor in human history.” This internal acknowledgment suggests that the entities responsible for building large language models were fully aware of the ethical and legal precariousness of their training datasets. By labeling the process as theft, these executives inadvertently validated the primary arguments long held by publishers: that AI utility is being built upon the unauthorized expropriation of human intellectual labor.

Bypassing Paywalls and Stripping Metadata

Beyond the rhetorical admissions, the unsealed material details the technical mechanics of these data practices. The documents suggest a systematic effort to bypass paywalls and scrape content at scale, a practice that directly contradicts the terms of service of many news organizations. Furthermore, the alleged deliberate stripping of copyright notices from training data indicates a calculated effort to sanitize datasets, potentially to evade the legal protections that safeguard original reporting and intellectual property.

The 'Existential Threat' to Journalism

Internal discussions at OpenAI reportedly centered on the concept of an “existential threat” posed to the publishing industry. Leadership within these firms recognized a “doom loop” dynamic, wherein the ingestion of high-quality journalism to train models would eventually lead to the erosion of the news organizations that produce such content. This recognition highlights a fundamental tension in the AI economy: the reliance on high-quality, human-curated data versus the economic degradation of the source of that data.

Broader Implications for AI Governance

The revelation of these internal documents is likely to serve as a pivotal moment in the legal landscape of generative AI. As courts grapple with the definition of 'fair use' in the context of machine learning, these admissions of internal dissent and awareness of harm may weaken the defense strategies employed by tech companies. This case will likely set a global precedent for how intellectual property rights are balanced against the rapid, often unchecked, advancement of artificial intelligence.

Future Trends in Data Ethics

Looking forward, the industry faces an inevitable shift toward more transparent and consensual data practices. The fallout from these disclosures will likely accelerate calls for standardized licensing models for AI training data. As legal scrutiny intensifies, tech companies may be forced to abandon 'scraping-first' methodologies in favor of partnerships that provide compensation to content creators, ultimately redefining the symbiotic, yet currently adversarial, relationship between AI developers and the media sector.

Verification Required?

Read the full report from the primary source

Go to Ars Technica - All content