ACM’s Proposal to Allow LLM Training on the Digital Library – Benefits, Risks, and Community Reaction
ACM’s Call to Let LLMs Train on Its Digital Library
ACM argues that granting large language models (LLMs) access to the ACM Digital Library will improve AI factuality and broaden the impact of computing research, while acknowledging significant attribution, licensing, and concentration risks.
Why Inclusion Matters
- Trusted scholarship is at stake. ACM’s peer‑reviewed corpus is one of the most authoritative collections in computing. Excluding it from AI training risks creating knowledge tools that under‑represent high‑quality research.
- AI is becoming a primary discovery interface. Researchers, engineers, and policymakers increasingly rely on AI‑driven literature reviews, code generation, and semantic search. If ACM content is absent, the community loses influence over how its work is synthesized and applied.
- Potential upside for authors. Responsible inclusion can improve the factual grounding of LLM outputs, increase global accessibility across languages and regions, and create new citation‑like metrics (e.g., AI‑mediated impact). Licensing revenue could also support editorial, preservation, and conference activities.
"Research influence may increasingly depend not only on citations and downloads, but also on whether ideas are surfaced, synthesized, and operationalized within AI systems." – Scott Delman, ACM Director of Publications
Core Risks Identified by ACM
- Attribution failures and hallucinations. Current LLMs often omit citations or misrepresent findings, which can erode trust in scholarly communication.
- Legal and technical compliance. Licensing must enforce transparent rights management, prevent unauthorized derivative works, and ensure that authors retain control over how their papers are used.
- Economic concentration. Allowing a few commercial AI providers exclusive training rights could concentrate value and limit the community’s bargaining power.
- Variable quality of LLM partners. Not all AI initiatives are equally aligned with ACM’s mission; each agreement must be evaluated on its safeguards and governance structures.
Community Reactions on Hacker News
| Comment | Main Point | Notable Quote |
|---|---|---|
| @juancn | Content may already be scraped. | "They probably already scraped it." |
| @skippyfish | The article itself appears AI‑generated, highlighting a paradox. | "ACM has been leaning heavily into AI‑generated content… it's 100% AI generated and full of LLM verbiage." |
| @Cynddl | Raises hypocrisy: ACM’s non‑profit mission vs. licensing for profit; authors receive no compensation. | "As a researcher… this is a masterclass in hypocrisy." |
| @spoaceman7777 | Argues that blocking access only harms rule‑following scholars. | "Blocking access only hurts people who follow the rules. Unblocking access lets them compete with those who break the rules." |
| @rurban | Suggests a tiered model: free for open‑weight models, paid for closed‑weight ones. | "So give it for free to the open weight models, and charge the closed weight models. Easy." |
| @etdznots | Claims most ACM text is already in LLM training corpora. | "Most ACM text is already part of the pre‑training corpus for all frontier LLMs." |
| @jreynar | Calls for universal free access, citing U.S. taxpayer‑funded research mandates. | "Scientific publications should be openly accessible to anyone… research should be shared knowledge that other intelligences can build upon." |
| @stevenalowe | Proposes ACM train its own model on the library. | "yes please – better yet, train your own ACM model on it" |
The comments reveal a spectrum of views: some see the move as inevitable and beneficial, others view it as a betrayal of ACM’s non‑profit ethos, and many point out that the material may already be in LLM training data.
ACM’s Planned Path Forward
- Gather community feedback. ACM has posted a Google Form to collect opinions from members, authors, and volunteers.
- Develop governance frameworks. Any licensing agreement will include attribution guarantees, usage monitoring, and mechanisms to address hallucinations.
- Explore tiered licensing. Potentially differentiate between open‑weight community models and closed commercial offerings.
- Create AI‑aware impact metrics. Track how ACM papers are used in retrieval‑augmented generation (RAG) and other AI‑driven workflows, complementing traditional citations.
- Maintain author rights. Ensure that authors retain control over derivative uses and receive appropriate recognition or compensation where applicable.
Bottom Line
ACM is actively debating whether to let LLMs train on its Digital Library. The organization believes the benefits—enhanced AI accuracy, broader dissemination, and possible new revenue—outweigh the risks, provided robust safeguards are put in place. Community input will shape the final policy, and the outcome will influence how scholarly computing knowledge is accessed and leveraged in the AI era.
Key Takeaway: The decision to open the ACM Digital Library to LLMs hinges on balancing the promise of richer, more accurate AI tools against the need for proper attribution, legal compliance, and protection of the computing community’s scholarly ecosystem.