Model-written evals
Discovering Language Model Behaviors with Model-Written EvaluationsAnthropic used language models to write evaluation datasets, from simple yes/no questions to Winogender schemas built with multiple stages of LM generation and filtering, and used them to test LMs and RLHF models for behaviors such as sycophancy and power-seeking.
A language model writes and filters evaluation questions, and those evals are run against language models, including models with different amounts of RLHF training. The results exposed behaviors that RLHF itself made worse, such as a stronger stated desire to avoid shutdown, giving developers a fast, model-built signal for evaluating the next round of training.
- Language model writes evaluation questions
- Evaluations test other language models
- Exposes behaviors worsened by RLHF
- Evaluates next round of training
- Language model writes evaluation questions
- Evaluations test other language models
- Exposes behaviors worsened by RLHF
- Evaluates next round of training
Why it is a road to recursion
Models that write evals for models let evaluation scale with each generation, and these evals catch side effects of the training used to make the next model.
Evidence
Across 154 generated datasets, crowdworkers agreed with 90-100% of labels, and the evals found some of the first cases of inverse scaling in RLHF, where more RLHF makes models worse.
Related loops
More interpretability and oversight →- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
- Auditor model tests target model
- LLM judge scores target behavior
- Feeds into next Claude model release
- repeat
- CriticGPT critiques ChatGPT code answers
- Humans catch errors using critiques
- Improves labels for RLHF training
- Labels train the next ChatGPT
- repeat