OpenAI Executives Acknowledged Illegal Book Piracy and Job Threats in Authors Guild Lawsuit

Executives admitted illegal use of pirated books

The Authors Guild lawsuit documents show that OpenAI and Microsoft deliberately incorporated copyrighted books from the LibGen repository, a known source of pirated material, into their training data. Internal Slack messages confirm that senior staff discussed "excising" LibGen files in 2022 to hide the evidence. This demonstrates conscious violation of copyright law.

"Given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage. What would be involved in that?" – OpenAI VP of Research Bob McGrew, June 15, 2022 (Class Plaintiffs’ Memorandum of Law).

Executives feared "optics" rather than legal risk

OpenAI leadership expressed concern that the use of LibGen could generate negative publicity, especially on Hacker News. Dario Amodei described the dataset as "a bit sketchier," and researcher Sam McCandlish said he was "just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on Hacker News would be unfortunate."

"I was just worried about optics – i.e. 'openai uses copyrighted data from sketchy russian website' showing up on HN would be unfortunate." – Sam McCandlish (Class Plaintiffs’ Memorandum of Law).

The comment highlights that the primary concern was reputational, not legal compliance.

Executives predicted their models would replace writers

Multiple internal statements indicate that OpenAI staff believed their models could displace human authors. Policy Director Jack Clark warned in May 2020 that "our work on AI and Creativity is going to increasingly lead to us creating systems that substitute for the labor of [] people… The better we do on GPT‑X, the more worried genre fiction authors will become about us substituting for them on Amazon."

"Our work in this area will make people unemployed… we will likely ignore their concerns and release anyway." – Jack Clark (internal memo, May 2020).

Tarun Gogineni, hired in 2022 to improve model writing quality, openly stated his research mission was to have GPT write the "last two books of [George R. R. Martin]'s A Song of Ice and Fire" and that he would "rest easy knowing that even if GRRM dies early, GPT‑5 will autocomplete his series."

"The world would soon experience ‘the death of the reader’ as ‘machines create slop for more machines.’" – Gogineni, internal notes (Class Plaintiffs’ Corrected Rule 56.1 Statement).

Microsoft was aware from the start

Microsoft executives were briefed on the LibGen usage as early as April 2019 when Sam Altman and Dario Amodei presented an early GPT‑3 prototype to Bill Gates and CTO Kevin Scott. The briefing included explicit disclosure of the LibGen source.

"Microsoft knew about OpenAI’s use of LibGen as early as April 2019 when Sam Altman and Dario Amodei presented an early version of GPT‑3 to Bill Gates and disclosed the use of LibGen." – Authors Guild filing.

Legal and cultural stakes

The plaintiffs argue that the defendants’ conduct constitutes willful infringement, which could dramatically increase statutory damages under U.S. copyright law. They also claim the conduct threatens the cultural ecosystem by substituting human creativity with AI‑generated text, which they describe as an "existential threat to those who write and publish books."

"OpenAI’s GPT models pose an existential threat to those who write and publish books." – Authors Guild CEO Mary Rasenberger.

Community reaction on Hacker News

Comments on the Hacker News discussion reflect a range of perspectives:

  • Some users dismissed the threat to jobs as typical technological disruption, comparing it to the impact of automobiles on horse‑drawn carriages.
  • Others criticized the framing of the lawsuit as a lobbying effort, questioning the neutrality of the Authors Guild press release.
  • A few highlighted the specific evidence of LibGen usage, noting that the dataset contains overwhelmingly copyrighted textbooks, undermining any "public‑domain" defense.
  • Several commenters expressed frustration that OpenAI seemed more concerned about negative coverage on Hacker News than about the legality of its data practices.

"Dario Amodei … was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on Hacker News would be unfortunate." – quoted in the post (Hacker News comment by papergirl).

What the filings mean for the AI industry

If the court finds that OpenAI and Microsoft acted willfully, damages could reach billions of dollars, setting a precedent that large‑scale AI training on unlicensed copyrighted works is illegal. The case also forces the industry to confront the ethical implications of building models that could replace creative professionals.

"We expect more briefing over the next couple months and a hearing in early 2027." – Authors Guild filing.

The outcome will likely shape future data‑collection policies, licensing negotiations, and possibly spur legislative action on AI‑trained‑on‑copyrighted‑material.


All quotations are taken directly from the Class Plaintiffs’ Memorandum of Law (Docket 1982) and the Class Plaintiffs’ Corrected Rule 56.1 Statement of Undisputed Material Facts (Docket 1987) linked in the Authors Guild press release.

Sources

Related