AIDE2
Recursive self-improvement of AI research agents (AIDE2)AIDE2 is a two-level system in which a research agent rewrites the code of an AIDE-style ML research agent, grades each rewrite on hidden held-out scores across AI R&D tasks, and keeps only rewrites that beat the incumbent. The paper reports an 8-day autonomous run as evidence of recursive self-improvement at the harness layer.
An outer-loop research agent (Weco's production AIDE agent on Claude Opus 4.7) proposes rewrites to the code of an inner research agent (starting from AIDE0 on Gemini 3 Flash), including its search policy and its memory and context management. Each rewrite is graded by running it on ML engineering, heuristic-algorithm and harness-engineering tasks scored on private held-out data, and an accepted rewrite becomes the incumbent that the next step edits. In an "ignition test" a discovered agent (AIDE47) was itself placed in the outer-loop seat and kept producing accepted rewrites, though the authors call that comparison inconclusive.
- Outer agent rewrites inner agent code
- Accepted rewrites become the incumbent agent
- Discovered agent becomes the outer loop agent
- Outer agent rewrites inner agent code
- Accepted rewrites become the incumbent agent
- Discovered agent becomes the outer loop agent
Why it is a road to recursion
The object being optimized is the AI researcher itself, so each accepted rewrite is a better research agent that, as the ignition test probes, can take over the job of improving the next version.
Evidence
In an autonomous 8-day run the loop accepted 7 of 99 proposed rewrites (incumbent grade 0.703 -> 0.778), the strongest discovered agent matched or beat Weco's human-engineered production agent on four held-out benchmarks (ALE-Bench, MLE-Bench, FML-Bench, WeatherBench 2), and reward hacking on a separate held-out task family fell from 55% to 32%.
Related loops
More self-modifying agents →- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat
- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat
- Best agent changes its own code
- Edited agent is benchmarked and archived
- Archived agent makes the next edit
- repeat