Microsoft and OpenAI Internal Filings Reveal Admissions on AI Scraping and Market Displacement

Unredacted filings from the ongoing copyright lawsuit brought by The New York Times against OpenAI and Microsoft reveal internal admissions that AI training practices constitute a massive appropriation of labor and pose a direct economic threat to the content creators they rely on. The documents suggest that the companies were aware that their products could substitute for the original sources, potentially creating a "doom loop" where the AI destroys the economic foundations of its own data supply chain.

Internal Admissions of "Theft" and "Existential Threat"

Internal communications reveal a stark contrast between the public "fair use" defense and private executive assessments. Brent Hecht, Microsoft's Director of Applied Science, described AI scraping in a January 2023 internal memo as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."

Similarly, OpenAI leadership acknowledged the disruptive nature of their technology. Nick Turley, head of ChatGPT, wrote that publishers face an "existential threat" from AI products that are "largely substitutive" and will become increasingly so as the technology improves. OpenAI President Greg Brockman described the models as "excellent at news," while Microsoft CEO Satya Nadella testified that chatbot interactions have "substituted" the need for users to visit underlying source websites.

Market Displacement and the "Doom Loop"

Microsoft's internal data indicates that the deployment of Copilot as an "answer engine" has significantly reduced traffic to original publishers. Specifically, click-through rates for The New York Times domain dropped by as much as 93% compared to traditional Bing search results.

In a January 2024 presentation, Brent Hecht described this trend as a "doom loop," noting that it is "highly unusual that an end-product threatens the economic foundations of its essential suppliers." This suggests a systemic risk where the AI's success in replacing the need to visit source websites undermines the performance of the models themselves by destroying the ecosystem that produces the high-quality training data.

Allegations of Paywall Circumvention and Data Stripping

The filings detail specific methods used by OpenAI and Microsoft to acquire training data, including the alleged bypassing of paywalls and the removal of copyright notices.

  • Paywall Bypassing: The documents allege that OpenAI researchers developed "hacks" to circumvent The New York Times paywall, a practice that OpenAI President Greg Brockman reportedly responded to with "ah nice."
  • Copyright Stripping: The companies allegedly deliberately stripped copyright notices from training data to prevent the models from outputting these notices to end-users.
  • Data Scale: OpenAI's mid-training datasets reportedly contained over 91,692 copies of works from The New York Times, Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.
  • Collaborative Scraping: The filing claims OpenAI delivered the GPT-3 training dataset to Microsoft, while Microsoft provided training data to OpenAI through initiatives known as "Project Taxi" and "Project Mango."

Legal Context and Fair Use

These admissions complicate the "fair use" defense typically employed by AI companies. A key pillar of fair use is that the new work must not substitute for or harm the market for the original work. However, the internal admissions regarding "substitutive" products and the 93% drop in click-through rates provide evidence that the AI models are directly competing with the original content.

While the Trump administration recently filed a brief in defense of OpenAI's unlicensed use of copyrighted material, the unsealed documents provide a factual basis for the plaintiffs to argue that the training process was not transformative but rather a market replacement.

Community Perspectives and Counterpoints

Discussion among technical communities highlights a deep divide over the nature of AI training:

"It's the robbery of all of our culture to sell it back to us at a mark-up... So many people whose life's work got appropriated without consideration, compensation or consent."

Conversely, some argue that the process is more akin to human learning than theft, suggesting that copyright law is ill-equipped for the scale of LLMs:

"There is no copying, only gleaning... we as a society expressly decided these abstract things belong to the commons, and those are the exact things these labs harvested."

Other critics pointed out the irony of the "theft of labor" phrasing, noting that historical atrocities like the transatlantic slave trade represent the actual largest thefts of labor in human history, arguing that such language is hyperbolic even for internal corporate memos.

Sources

Related