Darwin Gödel Machine
Darwin Godel Machine: Open-Ended Evolution of Self-Improving AgentsA system that grows an archive of coding agents, each of which rewrites its own Python codebase, and keeps variants that do well on coding benchmarks. Parents are sampled from the whole archive rather than only the current best agent.
A coding agent running on a frozen foundation model reads its own benchmark logs and edits its own code, for example adding better file-editing tools, long-context management and peer-review steps. Each child agent is scored on SWE-bench or Polyglot and added to the archive, and later self-modifications are made by agents drawn from that archive, so a gain in coding ability is also a gain in the agent that writes the next change.
- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
Why it is a road to recursion
The skill being improved (editing code) is the same skill used to make the next improvement, so benchmark gains compound into better self-modification.
Evidence
The DGM automatically raised its coding agent's score on SWE-bench from 20.0% to 50.0% and on Polyglot from 14.2% to 30.7%.
Related loops
More self-modifying agents →- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat
- Outer agent rewrites inner agent code
- Accepted rewrites become the incumbent agent
- Discovered agent becomes the outer loop agent
- repeat
- Best agent changes its own code
- Edited agent is benchmarked and archived
- Archived agent makes the next edit
- repeat