Constitutional AI
Constitutional AI: Harmlessness from AI FeedbackA method for training a harmless, non-evasive assistant where the only human input on harmlessness is a short list of principles (a constitution). The model critiques and revises its own responses for supervised fine-tuning, then is trained with reinforcement learning from AI feedback (RLAIF).
In the supervised stage the language model critiques its own responses against the constitution, revises them, and is fine-tuned on the revisions. In the RL stage a model judges which of two responses sampled from the fine-tuned model is better, those AI preferences train a preference model, and that preference model is the reward signal for RL, replacing human harmlessness labels.
- Model critiques and revises its own responses
- Model is fine-tuned on the revisions
- Model judges responses to train preference model
- Preference model provides RL reward signal
- Model critiques and revises its own responses
- Model is fine-tuned on the revisions
- Model judges responses to train preference model
- Preference model provides RL reward signal
Why it is a road to recursion
Once AI feedback can replace human labels as the training reward, each more capable model can act as critic and preference source for the next, removing human labeling as the bottleneck of the loop.
Evidence
Trained without any human feedback labels for harms, the RL-CAI assistant was preferred by crowdworkers over models trained with previously collected human harmfulness labels, and was less harmful at a given level of helpfulness.
Related loops
More self-reward and self-play →- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat
- Model proposes code reasoning tasks
- Model solves them for accuracy reward
- Python executor verifies the answers
- Curriculum shifts as solver improves
- repeat
- Policy model runs Monte Carlo Tree Search
- Verified trajectories train next policy and PPM
- PPM guides search to produce better data
- Self-evolution improves both models
- repeat