Microsoft says virtually nobody was grabbing NYT articles through its chatbot
Source Entity
Lauren Feiner

Microsoft is defending itself against copyright lawsuits by claiming its Copilot AI rarely reproduces substantial portions of copyrighted text. The company argues that an analysis of millions of chat logs shows minimal overlap with original news content.
Microsoft Defends AI Training Methods in Copyright Legal Battle
Microsoft is currently navigating a high-stakes legal confrontation regarding the copyright implications of its Copilot artificial intelligence. At the heart of the dispute is the contention from major publishers, including The New York Times, that the AI model unlawfully utilizes protected content to generate responses. Microsoft has countered these claims by asserting that its system does not function as a repository for copyrighted works, but rather as an engine that synthesizes information, rarely reproducing substantive chunks of text that could serve as a market substitute for the original journalism.
The Discovery Process and Data Analysis
As part of the discovery phase in these ongoing lawsuits, Microsoft turned over 8.2 million Copilot chat logs to experts representing the news publishers. This massive dataset was curated specifically to include interactions where users employed keywords related to the plaintiffs' websites, effectively creating a 'worst-case scenario' for potential copyright infringement. Despite this targeted selection process, Microsoft maintains that the resulting data does not support the publishers' claims of widespread content regurgitation.
Quantifying the Overlap
According to Microsoft’s filings, the analysis of these 8.2 million logs revealed that only 59,545 instances contained at least 16 words in common with the news content utilized to ground the model. By framing the data in this way, Microsoft is attempting to demonstrate that the frequency of verbatim reproduction is statistically negligible. The company’s stance suggests that the AI is performing a transformative task rather than acting as a simple copy-paste mechanism.
Implications for AI and Intellectual Property
The outcome of this litigation will likely set a significant precedent for how generative AI models are permitted to interact with proprietary content. If the courts accept Microsoft's argument that occasional word overlap does not constitute infringement, it could provide a legal buffer for other tech companies currently facing similar challenges. Conversely, a ruling in favor of the publishers could force a fundamental shift in how AI training data is licensed and processed.
Future Trends in AI Content Governance
Looking ahead, the tension between AI developers and content creators is expected to accelerate the development of more robust attribution and licensing frameworks. As AI models become more integrated into search and information retrieval, the legal definition of 'fair use' in the context of machine learning will be tested repeatedly. Microsoft’s focus on the rarity of substantial reproduction highlights a strategy aimed at proving that its technology respects the economic value of original reporting, even as it leverages that data to enhance user experience.
Conclusion
While Microsoft maintains its position that the Copilot system is not a vehicle for copyright theft, the legal battle remains far from settled. The reliance on expert analysis of chat logs underscores the complexity of quantifying AI behavior. As the case progresses, the industry will be watching closely to see how the court balances the innovative potential of generative AI against the established rights of publishers to control and monetize their intellectual property.