Anthropic Responsible Scaling Policy (RSP) Framework

Anthropic has implemented a Responsible Scaling Policy (RSP) to manage the unpredictable emergence of dangerous AI capabilities. The policy establishes a framework where the deployment and further training of models are contingent upon the implementation of specific safety safeguards, triggered by the model's actual capabilities rather than its size or version.

AI Safety Levels (ASL) Framework

Anthropic uses a system of AI Safety Levels (ASL) to categorize models based on their risk profile. Each level follows an "if-then" structure: if a model exhibits certain capabilities, then specific safeguards must be implemented before further deployment or training.

ASL-1: Low Risk

ASL-1 represents models with little to no risk, such as specialized AI systems designed for tasks like playing chess.

ASL-2: Present-Day Risk

ASL-2 represents the current state of AI models. These models possess a wide range of present-day risks but do not yet exhibit capabilities that could lead to catastrophic outcomes in fields such as chemistry or biology.

Required Safeguards for ASL-2:

  • Model cards
  • External red-teaming
  • Strong security measures

ASL-3: Catastrophic Misuse Risk

ASL-3 is triggered when AI models become operationally useful for catastrophic misuse in Chemical, Biological, Radiological, and Nuclear (CBRN) areas.

Required Safeguards for ASL-3:

  • Security measures strong enough to prevent non-state actors from stealing model weights and force state actors to expend significant effort to do so.
  • A guarantee that deployed versions of the model never produce information that operationally increases CBRN risks, even when red-teamed by world experts.
  • A rigorous definition of ASL-4 requirements before ASL-3 is reached.

ASL-4: Autonomous and Global Security Risk

ASL-4 is triggered when AI systems become capable of near-human level autonomy or become the primary source of a serious global security threat, such as bioweapons.

Required Safeguards for ASL-4:

  • A detailed and precise understanding of the model's internal workings to make an "affirmative case" that the model is safe.

Operationalizing the RSP

To ensure the RSP is a functional directive rather than a theoretical document, Anthropic employs several organizational strategies:

  • Executive Involvement: The policy was developed with significant time investment from the CEO and co-founders to signal the importance of AI safety to the entire organization.
  • Integration into Roadmaps: RSP protocols are treated as product and research requirements. Missing RSP deadlines can halt the training of new models or the shipping of products.
  • Resource Allocation: Anthropic has increased hiring in security, trust and safety, red teaming, and interpretability to meet the requirements of ASL-3.

Accountability and Governance

Anthropic ensures compliance with the RSP through a multi-layered governance structure:

  • Board Oversight: The RSP is a formal directive of the board.
  • External Oversight: The board is accountable to the Long Term Benefit Trust, an external panel of experts with no financial stake in the company.
  • Compliance Mechanisms: The company is implementing a whistleblower policy and has appointed an officer responsible for RSP compliance and reporting to the Long Term Benefit Trust.

Relationship with Regulation

Anthropic views the RSP as a prototype for future government regulation rather than a substitute for it. The goal is for companies to refine these frameworks and for governments to adopt the best elements to create standardized testing, auditing regimes, and oversight mechanisms. Anthropic advocates for a "race to the top" where companies and countries collaborate to improve safety frameworks.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch