Alphabell.
Foundation · Self-reward and self-play

Constitutional AI

Constitutional AI: Harmlessness from AI FeedbackA method for training a harmless, non-evasive assistant where the only human input on harmlessness is a short list of principles (a constitution). The model critiques and revises its own responses for supervised fine-tuning, then is trained with reinforcement learning from AI feedback (RLAIF).

The loop

In the supervised stage the language model critiques its own responses against the constitution, revises them, and is fine-tuned on the revisions. In the RL stage a model judges which of two responses sampled from the fine-tuned model is better, those AI preferences train a preference model, and that preference model is the reward signal for RL, replacing human harmlessness labels.

The loop
Constitutional AI
Foundation · arXiv
  1. Model critiques and revises its own responses
  2. Model is fine-tuned on the revisions
  3. Model judges responses to train preference model
  4. Preference model provides RL reward signal
  1. Model critiques and revises its own responses
  2. Model is fine-tuned on the revisions
  3. Model judges responses to train preference model
  4. Preference model provides RL reward signal
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Once AI feedback can replace human labels as the training reward, each more capable model can act as critic and preference source for the next, removing human labeling as the bottleneck of the loop.

Evidence

Trained without any human feedback labels for harms, the RL-CAI assistant was preferred by crowdworkers over models trained with previously collected human harmfulness labels, and was less harmful at a given level of helpfulness.

rlaifai-feedbackalignmentself-critiqueanthropic