Petri
Petri: An open-source auditing tool to accelerate AI safety researchAn open-source framework in which an auditor agent runs multi-turn, tool-using conversations with a target model from natural-language seed instructions, and LLM judges score the transcripts on safety-relevant dimensions for human review. Anthropic donated it to Meridian Labs with version 3.0 in May 2026.
An auditor model plays out scenarios against a target model and an LLM judge scores the target's behavior, so models test models at scale. Anthropic says Petri has been part of the alignment assessment of every Claude model since Claude Sonnet 4.5, so these AI-run audits feed into the release of each next Claude model.
- Auditor model tests target model
- LLM judge scores target behavior
- Feeds into next Claude model release
- Auditor model tests target model
- LLM judge scores target behavior
- Feeds into next Claude model release
Why it is a road to recursion
It makes auditing a reusable AI-run pipeline that each model generation both runs and is run against, which is how oversight can keep pace with capability.
Evidence
In the launch pilot across 14 frontier models and 111 seed instructions, Claude Sonnet 4.5 had the lowest overall misaligned-behavior score, slightly ahead of GPT-5.
Related loops
More interpretability and oversight →- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
- CriticGPT critiques ChatGPT code answers
- Humans catch errors using critiques
- Improves labels for RLHF training
- Labels train the next ChatGPT
- repeat
- Agent runs interpretability experiments
- Agent finds spurious background cues
- Selection retrains classifier final layer
- Improves robustness of interpreted model
- repeat