Anthropic Analysis of US Executive Order, G7 Code of Conduct, and Bletchley Park Summit

Anthropic has identified three pivotal events in AI policy from late 2023: the issuance of a comprehensive US Executive Order on AI, the establishment of a G7 International Code of Conduct, and the UK-led AI Safety Summit at Bletchley Park. These events signal a transition toward global, coordinated efforts to evaluate and monitor frontier AI systems for safety, security, and trustworthiness.

US Executive Order on AI

The US Executive Order (EO) establishes a comprehensive framework to address AI risks—including privacy, fairness, bias, and catastrophic risks—while directing government agencies to harness AI for improved public services through the appointment of Chief AI Officers.

Anthropic highlights several key technical and structural directives within the EO:

  • NIST Empowerment: The EO directs the National Institute of Standards and Technology (NIST) to develop evaluations for model capability and safety, expand the Secure Software Development Framework to include frontier models, and create a generative AI companion for the Risk Management Framework.
  • NAIRR Pilot: The order launches a pilot for the National AI Research Resource (NAIRR), designed to provide academic researchers with the data and compute necessary for AI safety research and the development of beneficial applications.

G7 International Code of Conduct

The G7 International Code of Conduct for Organizations Developing Advanced AI Systems, developed via the Hiroshima Process, provides a set of responsible practices for identifying and mitigating risks throughout the AI lifecycle.

This Code builds upon previous voluntary White House commitments and focuses on:

  • Risk Mitigation: Implementing evaluations, information sharing, governance, security procedures, and transparency measures.
  • Regulatory Baseline: Anthropic views the Code as a potential baseline for international best practices and future domestic regulations.

Bletchley Park Summit and the Bletchley Declaration

The UK government's AI Safety Summit convened global experts and officials to address frontier AI concerns, resulting in the Bletchley Declaration. This statement, signed by 28 countries including China and developing nations, calls for multi-stakeholder action to manage AI risks and benefits.

Key outcomes of the summit include:

  • International Research Assessment: Signatories agreed to support an international assessment of existing research on frontier AI capabilities and risks to inform governments on the state of the science.
  • Responsible Scaling Policy: During the summit, Anthropic CEO Dario Amodei presented the company's Responsible Scaling Policy as a potential prototype for regulatory approaches, though not as a replacement for formal regulation.

Establishment of Government AI Safety Institutes

A primary outcome of recent policy shifts is the creation of specialized government bodies dedicated to the technical evaluation of AI models.

UK AI Safety Institute

The UK's Frontier AI Task Force has been reconstituted as the AI Safety Institute, the first significant government initiative focused on evaluating frontier AI risks. This institute employs technical experts who have already conducted private testing on frontier models from labs including Anthropic.

US AI Safety Initiative

The United States is launching a companion to the UK institute via a NIST consortium. This initiative will focus on:

  • Evaluation Methods: Developing methods to evaluate AI systems for safety and trustworthiness.
  • Frontier Threats Red Teaming: NIST will develop guidelines for red teaming to identify dangerous dual-use capabilities in AI models.

The Importance of Evaluation Science

Anthropic asserts that advancing evaluation science and establishing independent testing protocols are foundational requirements for sensible regulation. Robust measurement capabilities allow governments to monitor AI effectively and provide an objective framework that enables less-resourced companies to compete with larger labs on the same safety benchmarks.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch