OpenAI and Journalism: Response to The New York Times Lawsuit

OpenAI has publicly responded to a lawsuit from The New York Times, stating that the claims are without merit and asserting that training AI models on publicly available internet materials constitutes fair use. The company maintains that it supports journalism and is actively working to eliminate the "regurgitation" of training data in its model outputs.

AI Training as Fair Use and the Opt-Out Mechanism

OpenAI contends that training AI models using publicly available internet materials is permitted as fair use under long-standing precedents. The company argues this principle is essential for innovators and critical for maintaining US competitiveness in AI.

To support this position, OpenAI cites a wide range of supporters who submitted comments to the US Copyright Office, including:

  • Academics
  • Library associations
  • Civil society groups
  • Startups
  • Leading US companies
  • Creators and authors

OpenAI also notes that similar laws permitting the training of models on copyrighted content exist in the European Union, Japan, Singapore, and Israel. Despite its legal stance on fair use, OpenAI provides a simple opt-out process for publishers to prevent its tools from accessing their sites, a process that The New York Times adopted in August 2023.

Addressing Content "Regurgitation"

OpenAI defines "regurgitation"—the memorization of specific training data—as a rare failure of the learning process. The company states that models are designed to learn general concepts to apply them to new problems, rather than to memorize specific text.

According to OpenAI, regurgitation is more common when content appears multiple times across different public websites. To combat this, the company has implemented measures to limit inadvertent memorization. OpenAI also emphasizes that intentionally manipulating models to regurgitate content is a violation of its terms of use.

Collaboration with News Organizations

OpenAI aims to support a healthy news ecosystem through partnerships that focus on three primary objectives:

  1. Operational Support: Deploying products to assist reporters and editors with time-consuming tasks, such as translating stories and analyzing voluminous public records.
  2. Model Enhancement: Training models on additional historical, non-publicly available content.
  3. Reader Connectivity: Displaying real-time content with attribution in ChatGPT to provide new ways for publishers to connect with readers.

Existing partnerships include collaborations with the Associated Press, Axel Springer, the American Journalism Project, and NYU.

Response to The New York Times Allegations

OpenAI claims that The New York Times is not presenting the full story regarding their legal dispute. The company states that negotiations for a high-value partnership regarding real-time display with attribution in ChatGPT were progressing constructively until December 19, 2023.

Regarding the allegations of content regurgitation, OpenAI asserts that:

  • The New York Times repeatedly refused to share examples of regurgitation during negotiations, despite OpenAI's commitment to investigate.
  • The examples cited by The New York Times appear to be years-old articles that proliferated on third-party websites.
  • The New York Times likely used adversarial prompts, including lengthy excerpts of articles, to intentionally induce the model to regurgitate content, which is not typical user activity.

Sources