NYT lawsuit reveals Microsoft exec calls AI scraping the "largest theft of labor in human history" and OpenAI calls ChatGPT an "existential threat" to publishers
Bottom line
Legal filings in the New York Times’ copyright lawsuit disclose that Microsoft’s applied‑science director Brent Hecht called AI‑driven web‑scraping “the largest theft of labor in human history,” while OpenAI’s head of ChatGPT labeled the product an “existential threat” to publishers. These statements undermine the companies’ public fair‑use defenses and could shape future litigation over large‑language‑model training data.
Key revelations from the NYT brief
- Microsoft’s internal memo (Jan 2023) – Hecht wrote that “millions of people … will soon consider large models ‘hoovering up’ all their work” and described the practice as “the largest theft of labor in human history.”
- Another Microsoft document warned that content creators “did not intend for their work to be used in this fashion, nor are they compensated for its use.”
- Impact on the NYT – Internal data cited by the brief shows Copilot reduced NYT click‑through rates by up to 93 % compared with Bing search, illustrating a direct market impact.
- OpenAI internal communications – Nick Turley, head of ChatGPT, called the chatbot “an existential threat” to publishers, saying it is “largely substitutive” and will become more so as the model improves. An OpenAI engineer testified that “no matter how prominently we show the links, users won’t click.”
- “Doom loop” memo – Hecht warned that the AI ecosystem creates a feedback loop that harms both the web’s content supply chain and model performance.
“It is highly unusual that an end‑product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’” – Brent Hecht, internal Microsoft memo (quoted in NYT brief)
Legal context
- The NYT sued OpenAI and Microsoft in late 2023 for alleged copyright infringement over thousands of news articles used to train ChatGPT and Copilot.
- The lawsuit is still pending in 2026; the NYT is now seeking summary judgment based on the newly disclosed internal documents.
- AI firms typically rely on the fair‑use defense, which hinges on factors such as purpose, nature of the work, amount used, and market effect. The NYT brief argues that the internal memos demonstrate awareness of market harm, weakening the fair‑use claim.
- In contrast, a U.S. court previously ruled that Anthropic’s use of published material qualified as fair use, illustrating the legal uncertainty.
Corporate positions on scraping
- Microsoft CEO Satya Nadella (deposition) – Stated that “anything that is paywalled should be licensed … for grounding or training.” He added that, had Microsoft known OpenAI scraped paywalled content, it would have required a model retraining.
- OpenAI’s public stance – Argues that large‑scale web scraping is permissible under fair use for research and training, citing the statutory definition that includes “research” and “scholarship.”
- Industry reaction – Some commentators note that the admissions could jeopardize the broader AI‑industry narrative that data scraping is benign.
Community perspectives (Hacker News comments)
- Persistent scraping concerns – Users report that Google’s AI Overview feature reproduces large verbatim excerpts from news articles, echoing the NYT’s complaint that AI systems embed source text without attribution.
- Plagiarism fears – A commenter described an LLM reproducing a novel term the user created, without citation, calling it “speed‑running a widespread corruption of truth.”
- Skepticism about corporate accountability – Others doubt that Microsoft would actually force OpenAI to retrain models, viewing the statement as “nice” but ineffective.
- Broader IP debate – Some participants argue that existing IP law is ill‑suited to the scale of AI training data, suggesting that current doctrines may be rendered obsolete.
Why this matters
- Litigation risk – The internal admissions provide concrete evidence that executives recognize direct economic damage to content creators, which could sway judges against a blanket fair‑use defense.
- Policy implications – If courts side with the NYT, future AI training may require explicit licensing of copyrighted material, reshaping data‑collection pipelines for large‑language models.
- Publisher survival – An “existential threat” label from OpenAI underscores the urgency for news organizations to develop mitigation strategies, such as paywall enforcement or AI‑generated content detection.
- Industry credibility – Public statements acknowledging market harm contrast sharply with the industry’s narrative of benign data use, potentially eroding trust among stakeholders and regulators.
Takeaways for developers and policymakers
- Audit training data – Companies should maintain transparent records of source material and consider licensing agreements for paywalled content.
- Implement attribution mechanisms – Embedding source citations in model outputs can reduce plagiarism claims and align with fair‑use factors.
- Monitor market impact – Quantify how AI products affect traffic and revenue for original content providers; such metrics may become evidentiary in future cases.
- Engage with legislators – Proactive dialogue on AI‑specific copyright reforms could preempt costly litigation and clarify permissible data‑use practices.
The information above is drawn from the New York Times’ legal brief filed in its copyright lawsuit against OpenAI and Microsoft, as reported by Tom’s Hardware on September 18 2026, and from discussion on Hacker News.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch