Anthropic Donates Petri Alignment Tool to Meridian Labs
Anthropic has transferred the development of Petri, its open-source alignment toolbox, to the nonprofit Meridian Labs and released version 3.0 of the tool. This move ensures that the alignment testing framework remains independent of any single AI laboratory, promoting neutral and credible evaluations of large language models (LLMs) across the industry.
Petri 3.0 Technical Enhancements
Petri 3.0 introduces three primary architectural and functional improvements designed to increase the accuracy and utility of alignment testing:
Increased Adaptability
Petri 3.0 features major architectural changes that decouple the auditor model from the target model. This separation allows users to tweak and modify each component independently, making the tool more adaptable to a wider variety of use cases.
Enhanced Realism via "Dish"
To prevent models from detecting they are being tested—which can lead to behavior that does not reflect general performance—Anthropic introduced an add-on called "Dish." Dish increases the realism of test environments by utilizing the model's actual system prompt and the real "scaffold" (the software wrapping the model to assist in goal achievement) that would be used in genuine deployments.
Deeper Behavioral Assessment through Bloom Integration
Petri is now integrated with Bloom, another open-source alignment tool from Anthropic. While Petri provides wide-ranging assessments, the integration with Bloom allows researchers to perform more in-depth evaluations of specific, chosen behaviors.
Operational Framework and Use Cases
Petri functions by simulating alignment-relevant scenarios using a separate "auditor" model. A subsequent "judge" model then scores the resulting transcripts to identify misaligned behaviors. The tool is specifically designed to detect concerning tendencies such as:
- Deception
- Sycophancy
- Cooperation with harmful requests
Anthropic has integrated Petri into the alignment assessment process for every Claude model starting with Claude Sonnet 4.5. Additionally, the UK's AI Security Institute (AISI) has utilized Petri as a core component of its evaluations regarding a model's propensity to sabotage AI research.
Transition to Meridian Labs
Anthropic has handed over the development of Petri to Meridian Labs, an AI evaluation nonprofit. This transition follows a similar pattern to Anthropic's donation of the Model Context Protocol (MCP) to the Linux Foundation. By moving Petri to a nonprofit entity, the tool joins a growing open technology stack—including other tools such as Inspect and Scout—available to labs, independent researchers, and governments to ensure reliable and neutral testing of AI model behavior.
Sources
Related
- Dispatch
- Dispatch
- Project
- Dispatch
- Dispatch