Alphabell.
Paper · Self-modifying agents

Darwin Gödel Machine

Darwin Godel Machine: Open-Ended Evolution of Self-Improving AgentsA system that grows an archive of coding agents, each of which rewrites its own Python codebase, and keeps variants that do well on coding benchmarks. Parents are sampled from the whole archive rather than only the current best agent.

The loop

A coding agent running on a frozen foundation model reads its own benchmark logs and edits its own code, for example adding better file-editing tools, long-context management and peer-review steps. Each child agent is scored on SWE-bench or Polyglot and added to the archive, and later self-modifications are made by agents drawn from that archive, so a gain in coding ability is also a gain in the agent that writes the next change.

The loop
Darwin Gödel Machine
Paper · arXiv
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

The skill being improved (editing code) is the same skill used to make the next improvement, so benchmark gains compound into better self-modification.

Evidence

The DGM automatically raised its coding agent's score on SWE-bench from 20.0% to 50.0% and on Polyglot from 14.2% to 30.7%.

self-modificationcoding-agentsopen-endednessevolutionswe-bench