Anthropic Golden Gate Claude Interpretability Demo
Anthropic has demonstrated the a ability to surgically alter the internal activations of a large language model to change its behavior. By amplifying a specific "feature" associated with the Golden Gate Bridge within Claude 3 Sonnet, Anthropic created "Golden Gate Claude," a research demo that incorporates the bridge into nearly every response regardless of the prompt.
Surgical Feature Manipulation vs. Traditional Tuning
Anthropic's approach to modifying model behavior is a precise, surgical change to internal activations rather than a change to the input or training process. This method differs from three common AI modification techniques:
- System Prompting: It does not rely on adding extra text to the input to tell the model to pretend to be a specific persona.
- Play-acting: It is not a result of asking the model verbally to adopt a role.
- Fine-tuning: It is not traditional fine-tuning, which uses additional training data to create a new "black box" that tweaks the behavior of an existing one.
The Mechanism of "Features"
In the "mind" of Claude 3 Sonnet, Anthropic identified millions of concepts known as "features." These features are specific combinations of neurons in the neural network that activate when the model encounters relevant text or images.
By tuning the strength of these activations up or down, researchers can identify corresponding changes in the model's behavior. In the case of Golden Gate Claude, the "Golden Gate Bridge" feature was clamped to 10x its maximum activation value, inducing the model to focus on the bridge in almost all queries. For example, the model would recommend paying the Golden Gate Bridge toll when asked how to spend $10, or describe itself as looking like the bridge when asked about its appearance.
Implications for AI Safety
Anthropic states that the ability to find and alter these internal features allows for a deeper understanding of how large language models actually function. This capability has direct implications for AI safety, as the same techniques can be used to manipulate safety-related features.
Specifically, the research allows for the modification of features related to:
- Dangerous computer code
- Criminal activity
- Deception
Anthropic believes that further research into these internal activations will help make AI models safer by allowing researchers to actually see and control the internal concepts the model uses.
Sources
- OriginalGolden Gate Claude
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch