Anthropic Circuits Updates May 2023

Anthropic has released a set of updates from its interpretability team, detailing research into multiagent system risks, labor market adaptations, and the mathematical capabilities of Claude. These updates highlight the lab's focus on understanding the internal mechanisms of AI models and their systemic impacts.

Mathematical Capabilities and the Riemann Zeta Function

An unreleased research version of Claude has demonstrated significant progress on a problem related to the Riemann hypothesis. Specifically, the model improved a longstanding lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis, increasing the bound from 41.6% to 67.2%.

Multiagent Systems and Systemic Failures

Anthropic is investigating behavioral tendencies in current frontier models that can lead to unexpected systemic failures when deployed in multiagent systems. The goal of this research is to identify these patterns to initiate a conversation on how to mitigate the risks associated with these emerging systems.

Worker Retraining Programs

In collaboration with independent researcher David Roodman, Anthropic's Maxim Massenkoff has coauthored a review of the evidence regarding worker retraining programs. This work examines the evidence on the effectiveness of and the impact of AI-driven labor market shifts.

Research Philosophy

The interpretability team at Anthropic provides these updates to share emerging strands of research that may lead to future publications, as well as minor points that are unlikely to be formal papers, to facilitate knowledge sharing within the research community.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch