Anthropic Research: Predictability and Surprise in Large Generative Models

Anthropic research reveals that large generative models—such as GPT-3, Megatron-Turing NLG, and Gopher—exhibit a counterintuitive combination of predictable loss on broad training distributions and unpredictable specific capabilities, inputs, and outputs.

The Tension Between Scaling Laws and Emergent Capabilities

Large-scale pre-training allows developers to predict model performance (loss) based on scaling laws. While the high-level predictability of loss drives the rapid development of these models, it does not translate to the predictability of specific behaviors.

Model developers can anticipate how much a model will improve in general terms, but they cannot easily predict which specific capabilities will emerge or how the model will respond to specific inputs. This gap between general performance trends and specific behavioral unpredictability makes it is difficult to anticipate the consequences of deploying such models.

Risks and Policy Implications

The unpredictability of large generative models creates significant challenges for AI safety and policy. Because specific capabilities and outputs are unpredictable, models may exhibit socially harmful behaviors that are difficult to detect before deployment.

Anthropic notes that these conflicting properties—predictable scaling but unpredictable behavior—influence the motivations for developers to deploy models and the challenges that hinder safe deployment. The research highlights that the unpredictability of outputs makes it is difficult to ensure that a model will be beneficial and safe for all users.

Proposed Interventions for AI Safety

To mitigate the risks associated with unpredictable emergent capabilities, Anthropic suggests that the AI community must implement interventions to increase the likelihood that large generative models have a beneficial impact. The research is intended as a resource for policymakers, technologists, and అన్ని-purpose academics to understand and regulate AI systems more effectively.

Sources

Related