OpenAI Investigation: Fake Russian Troll Error Message Hoax

OpenAI has identified a viral social media post claiming to expose a Russian troll account as a hoax. The incident is significant because it demonstrates how non-AI activity can be used to deceive the public about the actual use of AI models.

The Nature of the Hoax

The hoax involved a coordinated effort using an OpenAI account and an account on X (formerly Twitter). On June 18, the X account posted a comment that appeared to be a JSON error message from a Russian-speaking user who had run out of credits while attempting to generate content supportive of President Trump.

OpenAI's investigation confirmed that this specific error message was not generated by their models. The evidence supporting this conclusion includes:

  • The snippet of apparent JSON code was not valid JSON.
  • The message mis-referenced the name of the OpenAI model.
  • The content appears to have been manually created rather than AI-generated.

Actor Behavior and Model Usage

While the error message itself was fake, the actor behind the accounts did use OpenAI models in mid-June to generate short, adversarial comments to reply to other users on X. These comments covered a wide range of topics, including motorcycles, fantasy gaming, and flat-earth theories. The primary instruction given to the model was to be argumentative, regardless of the ideology involved.

OpenAI noted that some posts contained ableist terms used to denigrate others; however, these terms were not generated by the models but were added by another process before being posted to X.

Impact and Amplification

This incident represents a reversal of typical AI misuse cases. Instead of using AI to deceive people, the actor used non-AI activity to deceive people into believing AI was being used in a specific way.

Despite the original tweet having limited engagement (five reposts, 14 quotes, and three likes), the subsequent discussion and sharing of the tweet on LinkedIn, Reddit, and other platforms caused the content to go viral. OpenAI classifies this incident at the upper end of Category 3 on the Breakout Scale, nearly reaching Category 4 had mainstream media further amplified the event.

Outcome and Mitigation

Following the investigation, OpenAI banned the associated OpenAI account. Public reporting indicates that the X account involved in the hoax has also been suspended.

Sources