Alphabell.
Paper · Self-reward and self-play

Absolute Zero

Absolute Zero: Reinforced Self-play Reasoning with Zero DataA reinforcement learning paradigm in which one model proposes its own reasoning tasks and solves them, with a code executor validating the tasks and checking the answers, using no external data.

The loop

A single model proposes code reasoning tasks (deduction, abduction and induction) and is rewarded for tasks of moderate difficulty for itself, then solves them for an accuracy reward verified by a Python executor. Both roles are trained together, so the curriculum shifts as the solver improves.

The loop
Absolute Zero
Paper · arXiv
  1. Model proposes code reasoning tasks
  2. Model solves them for accuracy reward
  3. Python executor verifies the answers
  4. Curriculum shifts as solver improves
  1. Model proposes code reasoning tasks
  2. Model solves them for accuracy reward
  3. Python executor verifies the answers
  4. Curriculum shifts as solver improves
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

The model writes its own curriculum and is graded by an executor rather than by people, so training can continue past the supply of human-made tasks; the authors also report an 'uh-oh moment' in a Llama model's reasoning and call for safety-aware training.

Evidence

With no curated data, AZR on Qwen2.5-7B-Coder reaches a 50.4 combined code and math average (+10.2 over the base model), above every listed zero-setting reasoner trained on curated data (best 48.6).

self-playtask-proposalrlvrcode-executorzero-data