Confidence-Building Measures for Artificial Intelligence
OpenAI and the Berkeley Risk and Security Lab have published workshop proceedings detailing confidence-building measures (CBMs) designed to mitigate the risks foundation models pose to international security. These measures aim to reduce hostility, prevent the escalation of conflict, and build trust between global stakeholders to avoid accidents or unintentional conflict.
International Security Risks of Foundation Models
Foundation models introduce several pathways that could undermine state security. The primary risks identified include:
- Accidents and Inadvertent Escalation: The potential for AI-driven errors to lead to unintended conflict.
- Unintentional Conflict: Risks arising from the misinterpretation of AI capabilities or intentions.
- Weapon Proliferation: The proliferation of weapons facilitated by AI capabilities.
- Diplomatic Interference: The interference with human diplomacy by AI systems.
Proposed Confidence-Building Measures (CBMs)
Originating in the Cold War, confidence-building measures are flexible instruments used to reduce hostility and improve trust. The workshop participants identified six specific CBMs that apply directly to foundation models:
- Crisis Hotlines: Establishing direct communication channels to manage urgent security issues.
- Incident Sharing: The process of reporting and sharing information about AI-related security incidents.
- Model, Transparency, and System Cards: Utilizing standardized documentation to provide clarity on model capabilities and limitations.
- Content Provenance and Watermarks: Implementing technical markers to verify the origin of content and prevent misinformation.
- Collaborative Red Teaming and Table-top Exercises: Engaging in joint stress-testing of models to identify vulnerabilities.
- Dataset and Evaluation Sharing: Sharing the data and benchmarks used to evaluate model safety and performance.
Implementation and Stakeholder Involvement
Because the majority of foundation model developers are non-government entities, the implementation of these CBMs cannot rely solely on state-to-state agreements. Instead, these measures must involve a wider stakeholder community, including AI labs and government actors. These CBMs can be implemented either by the AI labs themselves or by relevant government agencies to ensure a stable and secure international environment.