Technology
Hacker News

Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

Source Entity

Hacker News

July 29, 2026
Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

AI firms are aggressively sourcing data by destroying rare books and scraping search results, sparking legal battles over copyright and fair use. Courts are currently navigating the tension between data-hungry AI development and the rights of content creators and platforms.

The War for Training Data: Destruction and Litigation

The landscape of artificial intelligence development has shifted toward an aggressive, high-stakes pursuit of training data. Recent reports indicate that AI companies are engaging in the industrial-scale destruction of rare books, physically removing spines and shredding original copies to facilitate high-speed scanning. Facilitated by services like ISBNdb, this practice prioritizes books published before 2022, which are valued for being untainted by AI-generated content. This physical erasure of cultural history for the sake of improving algorithmic outputs marks a concerning intersection between corporate greed and historical preservation.

The Legal Gray Area of 'Fair Use'

The justification for these actions often rests on a contentious interpretation of 'fair use.' A federal judge has reportedly ruled that destroying physical books to create digital copies is permissible under the logic that only one iteration of the content exists at any given time. This interpretation highlights a significant vulnerability in current copyright law, which struggles to reconcile the physical nature of historical artifacts with the digital-first requirements of large language models. The hiring of former Google Books leadership by firms like Anthropic underscores a strategic, long-term effort to aggregate global knowledge into proprietary AI datasets.

Hypocrisy in the Scraping Economy

Simultaneously, the broader tech ecosystem is embroiled in litigation regarding web scraping. Google recently attempted to utilize the Digital Millennium Copyright Act (DMCA) to block companies like SerpApi from scraping its own search results. This move has been widely criticized as hypocritical, given that Google’s own foundational business model was built on the systematic scraping of the open internet. By attempting to erect 'toll booths' around its search data, Google is effectively trying to pull up the ladder behind itself, while paradoxically arguing that its own anti-scraping measures are necessary to protect the rights of content creators.

Corporate Control vs. The Open Web

Recent court rulings have rejected Google’s attempts to use the DMCA to stifle competition, signaling a judicial reluctance to grant tech giants exclusive control over aggregated web data. However, the involvement of platforms like Reddit in similar efforts suggests a growing trend of 'walled garden' strategies. These platforms argue that scraping bots disrupt relationships with rights holders and threaten the integrity of their content, yet the underlying motivation is clearly to monetize data access in an era where high-quality human-generated text is becoming a scarce commodity.

Implications and Future Trends

As AI companies continue to treat physical archives and search databases as mere 'raw material,' we are likely to see an increase in legislative efforts to protect both rare literature and public digital information. The conflict between the 'data-is-free' ethos of early internet pioneers and the 'data-is-gold' reality of the AI era will define the next decade of intellectual property law. Unless clear boundaries are established, the cost of AI progress may be the irreparable loss of physical knowledge and the further balkanization of the internet into proprietary, gated communities.

Verification Required?

Read the full report from the primary source

Go to Hacker News