Alphabell.
Mini-project · Research automation

autoresearch-distillation

An open-source framework that trains an open-weight LLM agent with reinforcement learning (SDPO or GRPO) on live task-improvement loops, where each rollout edits a target file, runs the experiment on a GPU fleet and is rewarded by the measured metric. The flagship task is Karpathy's autoresearch: editing a GPT pretraining script to lower val_bpb in a 5-minute H100 run.

The loop

Qwen3-14B proposes edits to a small GPT's training script, each edit is trained for 5 minutes on a remote H100 and scored on validation bits per byte, and that score is the reward for an SDPO (or GRPO) update to Qwen3-14B's own weights. The model being trained is the ML researcher, so each update is meant to make it better at improving the next model's training run, and the trained checkpoint is re-run in the autoresearch loop to measure the gain.

The loop
autoresearch-distillation
Mini-project · GitHub
  1. Qwen3-14B edits GPT training script
  2. Edit is trained and scored
  3. Score rewards Qwen3-14B weights update
  4. Trained checkpoint re-run in autoresearch loop
  1. Qwen3-14B edits GPT training script
  2. Edit is trained and scored
  3. Score rewards Qwen3-14B weights update
  4. Trained checkpoint re-run in autoresearch loop
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It turns the outcomes of the autoresearch loop into training signal for the researcher's own weights, so a model that improves training scripts is itself trained to improve them better, the recursive step that prompt-only autoresearch lacks.

Evidence

In 50-turn autoresearch runs with feedback, the SDPO-trained Qwen3-14B reached 1.023 val_bpb (-3.1% from baseline) against 1.038 (-1.7%) for untrained Qwen3-14B and 1.027 (-2.8%) for Claude Opus 4.6, though single-turn Claude Haiku 4.5 posted the best absolute score (1.009, -4.4%) with high variance.

autoresearchreinforcement-learningsdpoopen-weightsresearch-agentverl