tech

Microsoft Exec Called AI Scraping ‘Largest Theft of Labor’ — Court Docs

Internal communications unsealed in the New York Times’ copyright lawsuit reveal Microsoft and OpenAI were acutely aware their data scraping could create a 'doom loop' for the web, with one executive calling the practice historic theft.

SignalEdge·September 19, 2026·4 min read
Symbolic image of a corporate executive in a server room, representing the 'largest theft of labor' in AI data scraping.

Key Takeaways

  • A Microsoft executive called OpenAI's web scraping for AI training the "largest theft of labor in human history," per unsealed court records.
  • Internal documents show both Microsoft and OpenAI feared their actions would create a "doom loop" for the web, harming publishers.
  • These revelations stem from the ongoing copyright infringement lawsuit filed by The New York Times against the two tech companies.
  • The documents suggest the companies understood the potentially damaging nature of their data collection methods even as they pursued them.

A Microsoft executive described OpenAI's web scraping for AI model training as the "largest theft of labor in human history," according to newly unsealed court documents. The explosive quote, reported by The Washington Post and Ars Technica, is part of a trove of internal communications revealed in the ongoing copyright lawsuit brought by The New York Times against Microsoft and its partner OpenAI.

The documents show that far from being oblivious to the consequences, executives at both companies were worried that training models like ChatGPT by scraping millions of news articles and other web content was not only legally questionable but existentially threatening to the online information ecosystem. This wasn't an accident; it was a known risk.

"Doom Loop" Fears and Acknowledged Theft

The internal discussions went beyond simple copyright concerns. According to The Verge, which also reviewed the filings, the companies' own documentation warned they were initiating a "doom loop." This scenario describes a feedback cycle where AI models ingest content from publishers, then produce summaries and answers that reduce traffic back to the original sources, eventually starving the publishers of revenue and causing them to fail. Without new, high-quality human-generated content, future AI models would have nothing to train on, degrading their own quality.

The consensus across all reports from The Verge, Ars Technica, Engadget, and The Washington Post is that these were not isolated comments. They reflect a pattern of internal anxiety. The characterization of web scraping as the "largest theft of labor in human history" by a Microsoft executive is particularly stark. It moves the conversation from a technical debate over 'fair use' to an admission of a fundamentally extractive process.

A Smoking Gun in the Copyright Battle

These revelations provide powerful ammunition for The New York Times. The publisher's lawsuit argues that its content was used without permission to build a competing product. Proving that Microsoft and OpenAI were not just aware of the potential harm but had explicitly labeled their own actions as 'theft' and worried about a 'doom loop' severely undermines any defense based on ignorance or good faith.

This suggests the companies made a calculated decision, weighing the immense commercial potential of their AI models against the known risks to the open web and the legal jeopardy they were inviting. The pattern indicates a strategy of moving fast, consolidating a market lead, and dealing with the legal and ethical fallout later. For the publishers suing them, this isn't just about compensation for past infringement; it's about whether the foundational business model of generative AI is built on a practice its own creators knew was damaging. The unsealed documents suggest they knew the answer all along.

SignalEdge Insight

  • What this means: Microsoft and OpenAI's 'fair use' legal defense is weakened by internal documents that frame their data collection as 'theft' and harmful.
  • Who benefits: The New York Times and other publishers suing AI companies, who now have evidence of intent and awareness of harm.
  • Who loses: Microsoft and OpenAI, whose public posture of building a better web is directly contradicted by their private concerns.
  • What to watch: How these documents influence settlement negotiations or a potential court ruling on the legality of using copyrighted web data for AI training.

Sources & References

Daily Newsletter

Stay ahead of the curve

Get the most important stories in tech, business, and finance delivered to your inbox every morning.

You might also like