Alphabell.
Project · Interpretability and oversight

Petri

Petri: An open-source auditing tool to accelerate AI safety researchAn open-source framework in which an auditor agent runs multi-turn, tool-using conversations with a target model from natural-language seed instructions, and LLM judges score the transcripts on safety-relevant dimensions for human review. Anthropic donated it to Meridian Labs with version 3.0 in May 2026.

The loop

An auditor model plays out scenarios against a target model and an LLM judge scores the target's behavior, so models test models at scale. Anthropic says Petri has been part of the alignment assessment of every Claude model since Claude Sonnet 4.5, so these AI-run audits feed into the release of each next Claude model.

The loop
Petri
Project · Anthropic
  1. Auditor model tests target model
  2. LLM judge scores target behavior
  3. Feeds into next Claude model release
  1. Auditor model tests target model
  2. LLM judge scores target behavior
  3. Feeds into next Claude model release
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It makes auditing a reusable AI-run pipeline that each model generation both runs and is run against, which is how oversight can keep pace with capability.

Evidence

In the launch pilot across 14 frontier models and 111 seed instructions, Claude Sonnet 4.5 had the lowest overall misaligned-behavior score, slightly ahead of GPT-5.

alignment-auditingopen-sourcellm-judgered-teamingclaude