Alphabell.
Paper · Self-reward and self-play

rStar-Math

rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep ThinkingA method that trains a small policy model and a process preference model for math from MCTS-generated, code-verified reasoning trajectories, then uses both for test-time search.

The loop

The policy model runs Monte Carlo Tree Search over code-augmented reasoning steps, producing verified trajectories and step-level Q-values that train the next policy and a process preference model (PPM). In later rounds the PPM guides the search to produce better data, and four rounds of this self-evolution (bootstrapped in round 1 with DeepSeek-Coder-V2-Instruct) improve both models.

The loop
rStar-Math
Paper · arXiv
  1. Policy model runs Monte Carlo Tree Search
  2. Verified trajectories train next policy and PPM
  3. PPM guides search to produce better data
  4. Self-evolution improves both models
  1. Policy model runs Monte Carlo Tree Search
  2. Verified trajectories train next policy and PPM
  3. PPM guides search to produce better data
  4. Self-evolution improves both models
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Policy and reward model improve each other round by round, each producing better training signal for the other, a co-evolution loop that needs no stronger teacher after the bootstrap round.

Evidence

After 4 rounds of self-evolution over 747k problems, MATH accuracy rose from 58.8% to 90.0% for Qwen2.5-Math-7B and from 41.4% to 86.4% for Phi3-mini-3.8B.

process-reward-modelmctsself-evolutionmath-reasoningsmall-models