CriticGPT
LLM Critics Help Catch LLM BugsOpenAI trained CriticGPT, a GPT-4-based model trained with RLHF to write critiques pointing out bugs in model-written code, to help human trainers evaluate model outputs more accurately during RLHF.
CriticGPT, itself trained with RLHF, critiques ChatGPT's answers so that human raters catch errors they would otherwise miss, improving the labels that train the next ChatGPT. It found hundreds of errors in ChatGPT training data that raters had marked flawless, and the authors describe critiques as the first step of recursive reward modeling.
- CriticGPT critiques ChatGPT code answers
- Humans catch errors using critiques
- Improves labels for RLHF training
- Labels train the next ChatGPT
- CriticGPT critiques ChatGPT code answers
- Humans catch errors using critiques
- Improves labels for RLHF training
- Labels train the next ChatGPT
Why it is a road to recursion
An AI critic trained with RLHF improves the human feedback used for RLHF, the base step of recursive reward modeling in which each model helps supervise the next.
Evidence
On code with naturally occurring LLM errors, model-written critiques were preferred over human critiques in 63% of cases, and models caught more bugs than human contractors paid for code review.
Related loops
More interpretability and oversight →- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
- Auditor model tests target model
- LLM judge scores target behavior
- Feeds into next Claude model release
- repeat
- Agent runs interpretability experiments
- Agent finds spurious background cues
- Selection retrains classifier final layer
- Improves robustness of interpreted model
- repeat