Alphabell.
Paper · Interpretability and oversight

CriticGPT

LLM Critics Help Catch LLM BugsOpenAI trained CriticGPT, a GPT-4-based model trained with RLHF to write critiques pointing out bugs in model-written code, to help human trainers evaluate model outputs more accurately during RLHF.

The loop

CriticGPT, itself trained with RLHF, critiques ChatGPT's answers so that human raters catch errors they would otherwise miss, improving the labels that train the next ChatGPT. It found hundreds of errors in ChatGPT training data that raters had marked flawless, and the authors describe critiques as the first step of recursive reward modeling.

The loop
CriticGPT
Paper · arXiv
  1. CriticGPT critiques ChatGPT code answers
  2. Humans catch errors using critiques
  3. Improves labels for RLHF training
  4. Labels train the next ChatGPT
  1. CriticGPT critiques ChatGPT code answers
  2. Humans catch errors using critiques
  3. Improves labels for RLHF training
  4. Labels train the next ChatGPT
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

An AI critic trained with RLHF improves the human feedback used for RLHF, the base step of recursive reward modeling in which each model helps supervise the next.

Evidence

On code with naturally occurring LLM errors, model-written critiques were preferred over human critiques in 63% of cases, and models caught more bugs than human contractors paid for code review.

scalable-oversightrlhfcritic-modelcode-reviewreward-modeling