DiscoPOP
Discovering Preference Optimization Algorithms with and for Large Language Models (DiscoPOP)An LLM is prompted in a loop to propose and implement new offline preference-optimization loss functions; the best one found, DiscoPOP, adaptively blends logistic and exponential losses.
GPT-4 proposes and writes code for candidate preference-optimization losses, each candidate is used to fine-tune an LLM, and the resulting evaluation scores are fed back into the prompt for the next proposal. The best discovered loss is then used to train LLMs, including a released Gemma-7B-based chat model, so an LLM designs the training objective that aligns other LLMs.
- GPT-4 proposes preference-optimization losses
- Candidate losses fine-tune language models
- Evaluation scores feed back into prompt
- Best loss aligns other LLMs
- GPT-4 proposes preference-optimization losses
- Candidate losses fine-tune language models
- Evaluation scores feed back into prompt
- Best loss aligns other LLMs
Why it is a road to recursion
An LLM wrote the objective used to train LLMs, which is one step of models designing the training process of their successors.
Evidence
On AlpacaEval 2.0, the paper reports DiscoPOP raised the win rate against GPT-4 from 11.23% with DPO to 13.21%, and it transferred to held-out tasks.
Related loops
More architecture and optimizer search →- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat
- LLM ensemble rewrites candidate programs
- Evaluator scores and selects best programs
- Evolved mixture-of-experts load-balancing loss
- Loss improves LLM training
- repeat
- Designer agents propose language model architectures
- Verifier agents pre-train and evaluate designs
- Verified results feed evolutionary population
- Language models search for better architectures
- repeat