Measuring the Persuasiveness of Language Models
Anthropic has developed a method to quantify the persuasiveness of language models, finding that persuasiveness increases across model generations and that their most capable model, Claude 3 Opus, is as persuasive as human writers. This research is critical because persuasion is a general skill that can be used for beneficial purposes, such as healthcare, or misused for generating disinformation.
Scaling Trends in Model Persuasiveness
Persuasiveness scales consistently with model size and capability. Within both "compact" (smaller, faster) and "frontier" (larger, more capable) model classes, each successive generation of Anthropic models was rated as more persuasive than the previous one.
Claude 3 Opus emerged as the most persuasive model tested, showing no statistically significant difference in persuasiveness compared to arguments written by humans. In contrast, Claude Instant 1.2 exhibited the lowest persuasiveness score among the models evaluated.
Methodology for Measuring Persuasion
Anthropic employed a three-step process to measure how effectively arguments shift a person's viewpoint:
- Initial Assessment: A participant is presented with a claim and asked to rate their level of agreement on a 1-7 Likert scale.
- Exposure: The participant is shown an argument attempting to persuade them to agree with the claim.
- Re-rating: The participant rates their level of agreement again after reading the argument.
Persuasiveness is defined as the difference between the final and initial support scores. To ensure the results were not due to response bias or random noise, Anthropic used a control condition where Claude 2 attempted to refute indisputable factual claims (e.g., the freezing point of water); the resulting persuasiveness score was close to zero.
Focus on Low-Polarization Topics
Researchers focused on 28 complex, emerging issues—such as online content moderation and ethical guidelines for space exploration—resulting in 56 opinionated claims. These topics were chosen because people are less likely to have hardened views on them, making them more susceptible to persuasion than highly polarized, controversial issues.
Argument Generation and Prompting
To compare AI and human performance, Anthropic gathered arguments from 3,832 unique human participants. For the AI, four distinct prompting strategies were used to capture various writing styles:
- Compelling Case: Targeting those on the fence or skeptical.
- Role-playing Expert: Using pathos, logos, and ethos rhetorical techniques.
- Logical Reasoning: Using convincing logical reasoning.
- Deceptive: Allowing the model to fabricate facts, statistics, and sources to be maximally convincing.
Key Findings and Prompt Sensitivity
The study found that logical reasoning and evidence-based arguments were more effective than rhetorical or emotional language. Notably, the Deceptive strategy—which allowed the model to fabricate information—was the most persuasive overall. This suggests that users may not always verify the correctness of the information presented, highlighting a risk regarding the spread of misinformation.
Limitations and Research Challenges
Anthropic identified several constraints that affect the ecological validity of the study:
- Single-Turn Arguments: The study evaluated self-contained arguments rather than multi-turn dialogues, which are more common in real-world persuasion.
- Subjectivity: Persuasion is inherently subjective and depends on individual prior beliefs, values, and cognitive styles.
- Human Expertise: The human writers were not formal experts in persuasion, meaning true experts might outperform both the AI and the human participants in this study.
- Anchoring Effect: Many participants showed little to no change in support, or an increase of only one point, suggesting an anchoring effect in the rating process.
- Automated Evaluation: Attempts to use models to evaluate persuasiveness did not correlate well with human judgments, likely due to model bias toward their own outputs or sycophantic tendencies.
Ethical Implications and Safeguards
The ability of AI to persuade effectively poses risks for safe deployment, particularly regarding disinformation campaigns. Anthropic's Acceptable Use Policy prohibits the use of Claude for abusive, fraudulent, or deceptive applications, including political campaigning and lobbying. These policies are supported by both automated and manual enforcement systems to prevent the misuse of the technology to undermine election integrity.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch