Anthropic Activating AI Safety Level 3 (ASL-3) Protections

Anthropic Activates ASL-3 Protections for Claude Opus 4

Anthropic has activated the AI Safety Level 3 (ASL-3) Deployment and Security Standards for the launch of Claude Opus 4. This move is a precautionary and provisional action taken because the company cannot definitively rule out ASL-3 risks—specifically regarding chemical, biological, radiological, and nuclear (CBRN) weapons—due to continued improvements in model capabilities.

The Responsible Scaling Policy (RSP) Framework

Anthropic's AI Safety Level (ASL) standards are governed by its Responsible Scaling Policy (RSP), which mandates higher levels of protection as models reach specific "Capability Thresholds."

  • Deployment Measures: These target specific categories of misuse, primarily focusing on reducing the risk that models assist in the development or acquisition of CBRN weapons.
  • Security Controls: These are designed to prevent the theft of model weights, which represent the core intelligence of the AI.

While all previous models were deployed under the baseline AI Safety Level 2 (ASL-2) Standard, Claude Opus 4 is the first to be deployed under ASL-3. Anthropic has ruled out the need for the ASL-4 Standard for Claude Opus 4, and has determined that Claude Sonnet 4 does not require ASL-3 protections.

ASL-3 Deployment Measures: Preventing CBRN Misuse

ASL-3 deployment measures are narrowly focused on preventing Claude Opus 4 from assisting with end-to-end CBRN workflows that provide additive value beyond what is possible without large language models. These measures specifically target "universal jailbreaks"—systematic attacks used to extract long chains of CBRN-related information.

To achieve this, Anthropic has implemented a three-part strategy:

1. Hardening the System Against Jailbreaks

Anthropic uses Constitutional Classifiers, which are real-time guards trained on synthetic data of harmful and harmless CBRN-related prompts and completions. These classifiers monitor inputs and outputs to block harmful CBRN information with moderate compute overhead.

2. Jailbreak Detection

Monitoring is bolstered by a bug bounty program specifically for stress-testing Constitutional Classifiers, alongside offline classification systems and threat intelligence partnerships to identify universal jailbreaks quickly.

3. Iterative Defense Improvement

Anthropic remediates jailbreaks by generating synthetic jailbreaks that mimic discovered attacks and using that data to train updated classifiers.

ASL-3 Security Controls: Protecting Model Weights

Security controls under ASL-3 are designed to defend against sophisticated non-state actors attempting to steal model weights. The framework includes over 100 different security controls combining prevention and detection.

Key security measures include:

  • Standard Best Practices: Two-party authorization for weight access, enhanced change management protocols, and binary allowlisting for endpoint software.
  • Egress Bandwidth Controls: A specialized control that restricts the flow of data out of secure computing environments. Because model weights are substantial in size, limiting outbound network traffic allows Anthropic to detect and block unusual bandwidth usage that may indicate an exfiltration attempt.

Implementation Rationale and Future Outlook

Anthropic implemented ASL-3 protections proactively to refine these defenses before they became mandatory. The company notes that dangerous capability evaluations are challenging and that as models approach thresholds of concern, determining their exact status takes longer.

If further evaluation concludes that Claude Opus 4 has not surpassed the relevant Capability Threshold, Anthropic may adjust or remove these ASL-3 protections. The company intends to continue iterating on these measures and collaborating with industry partners, government, and civil society to improve AI guarding methods.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch