OpenAI Disrupts Covert AI Influence Operations
OpenAI has disrupted five covert influence operations (IO) that attempted to use its AI models to manipulate public opinion and influence political outcomes. Despite the use of generative AI, none of these campaigns achieved a meaningful increase in audience engagement or reach, with all scoring a 2 or lower on the Brookings Breakout Scale.
Disruption of Five Covert Influence Operations
OpenAI identified and terminated accounts associated with five distinct threat actors who used AI for tasks such as generating political comments, translating articles, conducting open-source research, and debugging code.
Identified Threat Actors
- Bad Grammar (Russia): A previously unreported operation targeting Ukraine, Moldova, the Baltic States, and the United States. Actors used AI to debug Telegram bot code and generate short political comments in Russian and English.
- Doppelganger (Russia): This operation used AI to generate comments in English, French, German, Italian, and Polish for X and 9GAG, translate and edit articles, and convert news stories into Facebook posts.
- Spamouflage (China): This network used AI to research public social media activity, debug code for managing websites (including revealscum[.]com), and generate texts in Chinese, English, Japanese, and Korean for X, Medium, and Blogspot.
- International Union of Virtual Media (IUVM, Iran): This operation used AI to generate and translate long-form articles, headlines, and website tags for the site iuvmpress[.]co.
- Zero Zeno (Israel): Linked to a commercial company called STOIC, this operation used AI to generate articles and comments posted across Instagram, Facebook, X, and associated websites.
The content generated by these actors focused on global geopolitical issues, including the invasion of Ukraine, the conflict in Gaza, the Indian elections, and criticisms of the Chinese government.
Attacker Trends in AI Usage
Investigations into these operations reveal that AI is being used to enhance productivity rather than replace human operators. Key trends include:
- Increased Volume and Quality: AI allowed actors to produce text and images in higher volumes and with fewer linguistic errors than human operators could achieve alone.
- Hybrid Content Strategies: AI was not used exclusively; it was mixed with traditional formats, such as manually written texts and copied memes.
- Simulated Engagement: Some networks used AI to generate replies to their own posts to create a false appearance of engagement, though this did not result in authentic audience growth.
- Operational Productivity: AI was leveraged for non-content tasks, such as summarizing social media posts and debugging code.
Defensive Trends and Mitigations
OpenAI employs several defensive strategies to mitigate the impact of covert influence operations and improve detection speed.
Safety Systems and Design
OpenAI integrates safety systems into its model design to create friction for threat actors. This resulted in multiple instances where models refused to generate the requested deceptive content.
AI-Powered Investigations
By using internal AI-powered tools, OpenAI reduced the time required for investigations from weeks or months to just a few days. These tools are used for detection and analysis, similar to how GPT-4 is used for content moderation.
Distribution and Human Error
Because AI-generated material must be distributed via third-party platforms to reach an audience, the distribution phase remains a critical point of failure for attackers. Furthermore, OpenAI noted that threat actors remain prone to human error, such as accidentally publishing the model's refusal messages on their public websites and social media profiles.
Industry Collaboration
OpenAI shares detailed threat indicators with industry peers and relies on open-source research from the wider community to identify and disrupt these operations more effectively.