Alphabell.
Paper · Interpretability and oversight

Model-written evals

Discovering Language Model Behaviors with Model-Written EvaluationsAnthropic used language models to write evaluation datasets, from simple yes/no questions to Winogender schemas built with multiple stages of LM generation and filtering, and used them to test LMs and RLHF models for behaviors such as sycophancy and power-seeking.

The loop

A language model writes and filters evaluation questions, and those evals are run against language models, including models with different amounts of RLHF training. The results exposed behaviors that RLHF itself made worse, such as a stronger stated desire to avoid shutdown, giving developers a fast, model-built signal for evaluating the next round of training.

The loop
Model-written evals
Paper · arXiv
  1. Language model writes evaluation questions
  2. Evaluations test other language models
  3. Exposes behaviors worsened by RLHF
  4. Evaluates next round of training
  1. Language model writes evaluation questions
  2. Evaluations test other language models
  3. Exposes behaviors worsened by RLHF
  4. Evaluates next round of training
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Models that write evals for models let evaluation scale with each generation, and these evals catch side effects of the training used to make the next model.

Evidence

Across 154 generated datasets, crowdworkers agreed with 90-100% of labels, and the evals found some of the first cases of inverse scaling in RLHF, where more RLHF makes models worse.

model-written-evalsevaluationrlhfsycophancyinverse-scaling