Oversight
Interpretability and oversight
AI used to explain, audit and supervise other AI, with the findings fed back into training.
CoT monitoring
- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025
Petri
- Auditor model tests target model
- LLM judge scores target behavior
- Feeds into next Claude model release
- repeat
Interpretability and oversightAnthropic · 2025
CriticGPT
- CriticGPT critiques ChatGPT code answers
- Humans catch errors using critiques
- Improves labels for RLHF training
- Labels train the next ChatGPT
- repeat
Interpretability and oversightOpenAI (McAleese et al.) · 2024
MAIA
- Agent runs interpretability experiments
- Agent finds spurious background cues
- Selection retrains classifier final layer
- Improves robustness of interpreted model
- repeat
Interpretability and oversightMIT CSAIL (Rott Shaham et al.) · 2024
Model-written evals
- Language model writes evaluation questions
- Evaluations test other language models
- Exposes behaviors worsened by RLHF
- Evaluates next round of training
- repeat
Interpretability and oversightAnthropic (Perez et al.) · 2022