Anthropic Collaboration with US CAISI and UK AISI for AI Safeguards

Anthropic has established an ongoing partnership with the US Center for AI Standards and Innovation (CAISI) and the UK AI Security Institute (AISI) to identify and mitigate vulnerabilities in its AI models. This collaboration leverages government expertise in national security, cybersecurity, and threat modeling to improve the robustness of AI safeguards against sophisticated misuse.

Strengthening Constitutional Classifiers through Red-Teaming

Anthropic utilized CAISI and AISI to evaluate several iterations of its Constitutional Classifiers—the defense systems used to detect and prevent jailbreaks—on models including Claude Opus 4 and 4.1 prior to their deployment. This iterative stress-testing process uncovered several critical vulnerability classes:

  • Prompt Injection: Government red-teamers identified weaknesses where specific annotations, such as falsely claiming a human review had occurred, could bypass classifier detection.
  • Universal Jailbreaks: Testers developed sophisticated jailbreaks that encoded harmful interactions to evade standard detection. This led Anthropic to fundamentally restructure its safeguard architecture rather than applying a simple patch.
  • Cipher-based Attacks: The use of ciphers, character substitutions, and other obfuscation techniques was used to evade classifiers, leading to improvements in detection systems to recognize disguised content regardless of encoding.
  • Input and Output Obfuscation: Testers discovered universal jailbreaks that fragmented harmful strings into benign components within a larger context, resulting in targeted improvements to filtering mechanisms.
  • Automated Attack Refinement: The partners developed automated systems to progressively optimize attack strategies, producing effective universal jailbreaks that Anthropic is now using to improve its safeguards.

Enhancing Risk Methodology and Evaluation

Beyond specific technical vulnerabilities, the collaboration with CAISI and AISI has improved Anthropic's broader security approach. The government bodies provided an external perspective on deployment monitoring, rapid response capabilities, and evidence requirements, which helped pressure-test Anthropic's internal threat models and identify areas where more evidence was required.

Framework for Effective Government-Industry Collaboration

Anthropic identifies several key factors that enabled the effectiveness of this partnership with government research and standards bodies:

Comprehensive System Access

Providing deep access to systems allows for more sophisticated vulnerability discovery. Anthropic provided the following resources to government red-teamers:

  • Pre-deployment Prototypes: Access to safeguard prototypes allowed weaknesses to be identified before the systems went live.
  • Diverse System Configurations: Testers had access to models ranging from completely unprotected base models to those with full safeguards, allowing them to refine attacks progressively.
  • Internal Documentation: Transparency regarding safeguard architecture, documented vulnerabilities, and granular content policy information helped testers target high-value areas.
  • Real-time Data: Direct access to classifier scores enabled testers to refine attack strategies and conduct targeted exploratory research.

Iterative and Sustained Engagement

Sustained collaboration, including daily communication and frequent technical deep-dives, allows external teams to develop the deep system expertise necessary to uncover complex vulnerabilities that single evaluations cannot find.

Multi-Layered Defense Strategy

Anthropic combines specialized government evaluations with other security measures, such as public bug bounty programs. While bug bounties provide a high volume of diverse reports from a wide talent pool, government teams provide the deep technical knowledge required to uncover subtle, complex attack vectors.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch