Genesys
Language Modeling by Language Models (Genesys)Genesys is a multi-agent LLM system that simulates the research cycle for new language-model architectures, from ideation and literature search to code, pre-training and evaluation, on a genetic-programming backbone.
LLM designer agents propose, adversarially review and implement new LM architecture designs, and verifier agents pre-train and evaluate selected designs on a Ladder of Scales (14M to 350M parameters) with fewer models trained at each larger scale. Verified results feed the evolutionary population the designers draw from next, so language models search for better language-model architectures.
- Designer agents propose language model architectures
- Verifier agents pre-train and evaluate designs
- Verified results feed evolutionary population
- Language models search for better architectures
- Designer agents propose language model architectures
- Verifier agents pre-train and evaluate designs
- Verified results feed evolutionary population
- Language models search for better architectures
Why it is a road to recursion
It is an explicit setup of language models discovering language-model designs, the shape a recursive loop takes once the discovered designs train the next designers.
Evidence
Of 1,162 newly discovered designs (1,062 verified through pre-training), the best outperform GPT2, Mamba2 and others on 6 of 9 common benchmarks, and the genetic-programming backbone gave about an 86 percentage point improvement in successful design generation over direct prompting.
Related loops
More architecture and optimizer search →- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat
- LLM ensemble rewrites candidate programs
- Evaluator scores and selects best programs
- Evolved mixture-of-experts load-balancing loss
- Loss improves LLM training
- repeat
- GPT-4 proposes preference-optimization losses
- Candidate losses fine-tune language models
- Evaluation scores feed back into prompt
- Best loss aligns other LLMs
- repeat