Forecasting Misuse of Language Models for Disinformation Campaigns
OpenAI, in collaboration with Georgetown University’s Center for Security and Emerging Technology and the Stanford Internet Observatory, has released a report investigating the potential for large language models (LLMs) to be misused in disinformation campaigns. The research concludes that LLMs will likely transform online influence operations by reducing costs and increasing the scale and persuasiveness of deceptive content.
Impact of LLMs on Influence Operations
Large language models have the potential to alter the three primary components of influence operations: actors, behavior, and content.
Actors
LLMs may lower the financial and technical barriers to entry for running influence operations. This enables new types of actors to enter the field and provides a competitive advantage to "propagandists-for-hire" who can automate text production.
Behavior
The scale of influence operations can be increased through LLMs, making expensive tactics—such as the generation of highly personalized content—significantly cheaper. Additionally, LLMs enable new tactics, including the use of chatbots for real-time content generation.
Content
LLMs can produce messaging that is more persuasive or impactful than human propagandists, particularly when the propagandist lacks the cultural or linguistic knowledge of the target audience. Furthermore, LLMs make these operations harder to detect because they can generate unique content continuously, avoiding the repetitive copy-pasting patterns typically used to save time in manual campaigns.
Critical Unknowns in AI-Enabled Disinformation
Despite the identified risks, several variables remain unknown and require further research to determine the actual extent of the threat:
- Emergent Capabilities: Which well-intentioned commercial or research investments will inadvertently create new capabilities for influence operations?
- Actor Investment: Which specific actors will invest most heavily in LLM technology?
- Tool Availability: When will easy-to-use text generation tools become public, and will specialized models for influence be more effective than generic ones?
- Norms and Intentions: Will social or political norms develop to disincentivize the use of AI in influence operations?
A Four-Stage Framework for Mitigations
To address these threats, the report introduces a pipeline framework that identifies four stages where mitigations can be applied to disrupt an influence operation.
| Stage in the Pipeline | Illustrative Mitigations |
|---|---|
| 1. Model Construction | Building fact-sensitive models; spreading "radioactive data" to make models detectable; government restrictions on data collection or AI hardware access. |
| 2. Model Access | Implementing stricter usage restrictions; developing new norms for model release; closing security vulnerabilities. |
| 3. Content Dissemination | Coordinating between platforms and AI providers to identify AI content; requiring "proof of personhood" for posting; adopting digital provenance standards. |
| 4. Belief Formation | Engaging in media literacy campaigns; providing consumer-focused AI tools to help users identify misinformation. |
Evaluating Mitigation Desirability
The report emphasizes that the existence of a mitigation does not automatically make it desirable. Policymakers and developers are encouraged to evaluate potential interventions using four guiding questions:
- Technical Feasibility: Does the mitigation require significant changes to existing technical infrastructure?
- Social Feasibility: Is the mitigation actionable under current law and regulation, and are key actors incentivized to implement it?
- Downside Risk: What are the potential negative impacts of the mitigation?
- Impact: How effective is the proposed mitigation at actually reducing the threat?
Regardless of whether advanced models are kept private or restricted via API, the researchers anticipate that propagandists will utilize open-source alternatives or that nation-states will develop their own internal capabilities.