Alphabell.
Paper · Architecture and optimizer search

DiscoPOP

Discovering Preference Optimization Algorithms with and for Large Language Models (DiscoPOP)An LLM is prompted in a loop to propose and implement new offline preference-optimization loss functions; the best one found, DiscoPOP, adaptively blends logistic and exponential losses.

The loop

GPT-4 proposes and writes code for candidate preference-optimization losses, each candidate is used to fine-tune an LLM, and the resulting evaluation scores are fed back into the prompt for the next proposal. The best discovered loss is then used to train LLMs, including a released Gemma-7B-based chat model, so an LLM designs the training objective that aligns other LLMs.

The loop
DiscoPOP
Paper · arXiv
  1. GPT-4 proposes preference-optimization losses
  2. Candidate losses fine-tune language models
  3. Evaluation scores feed back into prompt
  4. Best loss aligns other LLMs
  1. GPT-4 proposes preference-optimization losses
  2. Candidate losses fine-tune language models
  3. Evaluation scores feed back into prompt
  4. Best loss aligns other LLMs
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

An LLM wrote the objective used to train LLMs, which is one step of models designing the training process of their successors.

Evidence

On AlpacaEval 2.0, the paper reports DiscoPOP raised the win rate against GPT-4 from 11.23% with DPO to 13.21%, and it transferred to held-out tasks.

preference-optimizationllm-discoveryloss-functionsrlhfsakana