Alphabell.
Paper · Self-reward and self-play

Self-Taught Evaluator

Self-Taught EvaluatorsA method for training an LLM-as-a-Judge with no human preference labels, using only synthetic contrasting response pairs and the judge's own filtered reasoning.

The loop

An LLM builds a worse answer for each instruction by answering a subtly modified version of it, which yields a preference pair with a known winner. The current judge samples reasoning traces and verdicts, only verdicts that pick the known winner are kept, the judge is fine-tuned on them, and the improved judge labels the next iteration's data.

The loop
Self-Taught Evaluator
Paper · arXiv
  1. LLM builds worse answers for preference pairs
  2. Judge samples reasoning traces and verdicts
  3. Judge is fine-tuned on correct verdicts
  4. Improved judge labels next iteration's data
  1. LLM builds worse answers for preference pairs
  2. Judge samples reasoning traces and verdicts
  3. Judge is fine-tuned on correct verdicts
  4. Improved judge labels next iteration's data
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Judges supply the reward signal for training other models, so an evaluator that improves itself without human labels removes a human bottleneck on the feedback that trains every model downstream.

Evidence

Without labeled preference data, the method raised Llama3-70B-Instruct from 75.4 to 88.3 on RewardBench (88.7 with majority vote), matching top reward models trained on labeled examples.

llm-as-judgereward-modelsynthetic-preferencesiterative-trainingevaluation