Self-Rewarding LMs
Self-Rewarding Language ModelsA training method in which one language model both generates responses and judges them through LLM-as-a-Judge prompting, using its own judgments for preference training.
The model generates several candidate responses to new prompts, scores them itself with an LLM-as-a-Judge prompt, and turns the highest and lowest scored into preference pairs for DPO. Each trained model becomes the generator and judge for the next iteration, and the paper reports that both instruction following and the quality of its self-rewards improve across iterations.
- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
Why it is a road to recursion
The reward model is not frozen but is the same model being trained, so improving the policy also improves the judge, which the authors argue is needed to get feedback beyond human level.
Evidence
Over three iterations on Llama 2 70B, the AlpacaEval 2.0 win rate rose from 9.94% to 15.38% to 20.44%, outperforming Claude 2, Gemini Pro and GPT-4 0613 on that leaderboard.
Related loops
More self-reward and self-play →- Model proposes code reasoning tasks
- Model solves them for accuracy reward
- Python executor verifies the answers
- Curriculum shifts as solver improves
- repeat
- Policy model runs Monte Carlo Tree Search
- Verified trajectories train next policy and PPM
- PPM guides search to produce better data
- Self-evolution improves both models
- repeat
- LLM builds worse answers for preference pairs
- Judge samples reasoning traces and verdicts
- Judge is fine-tuned on correct verdicts
- Improved judge labels next iteration's data
- repeat