Alphabell.
Paper · Self-reward and self-play

Self-Rewarding LMs

Self-Rewarding Language ModelsA training method in which one language model both generates responses and judges them through LLM-as-a-Judge prompting, using its own judgments for preference training.

The loop

The model generates several candidate responses to new prompts, scores them itself with an LLM-as-a-Judge prompt, and turns the highest and lowest scored into preference pairs for DPO. Each trained model becomes the generator and judge for the next iteration, and the paper reports that both instruction following and the quality of its self-rewards improve across iterations.

The loop
Self-Rewarding LMs
Paper · arXiv
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

The reward model is not frozen but is the same model being trained, so improving the policy also improves the judge, which the authors argue is needed to get feedback beyond human level.

Evidence

Over three iterations on Llama 2 70B, the AlpacaEval 2.0 win rate rose from 9.94% to 15.38% to 20.44%, outperforming Claude 2, Gemini Pro and GPT-4 0613 on that leaderboard.

llm-as-judgeself-rewarditerative-dpoalignmentpreference-learning