Paper · Self-reward and self-play
The Red Queen Gödel Machine
The Red Queen Gödel Machine: Co-Evolving Agents and Their EvaluatorsAn evolutionary framework that enables recursive self-improvement by co-evolving agents alongside the evaluators that guide their search.

The loop
A coder agent and a reviewer agent co-evolve under non-stationary utilities. The reviewer grades patches to guide the coder's search, while the coder's output helps the reviewer refine its grading rubric, making both better at their respective tasks.
The loop
The Red Queen Gödel Machine- Coder generates patches
- Reviewer grades patches to guide search
- Coder output helps reviewer refine rubric
- Reviewer provides better guidance to coder
- Coder generates patches
- Reviewer grades patches to guide search
- Coder output helps reviewer refine rubric
- Reviewer provides better guidance to coder
↻ The improved system does the next round, and the loop turns again.
Why it is a road to recursion
Co-evolving the evaluator alongside the agent removes the ceiling imposed by static reward functions, enabling open-ended recursive self-improvement.
Evidence
At low reasoning effort, the RQGM coder passes 82.1% of held-out tasks against the baseline's 75.0%.
co-evolutionagentic-codinglearned-evaluators
Related loops
More self-reward and self-play →
Self-Rewarding LMs
- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024

Absolute Zero
- Model proposes code reasoning tasks
- Model solves them for accuracy reward
- Python executor verifies the answers
- Curriculum shifts as solver improves
- repeat
Self-reward and self-playTsinghua, BIGAI, Penn State (Zhao et al.) · 2025

rStar-Math
- Policy model runs Monte Carlo Tree Search
- Verified trajectories train next policy and PPM
- PPM guides search to produce better data
- Self-evolution improves both models
- repeat
Self-reward and self-playMicrosoft Research Asia · 2025