Alphabell.
Dataset · Training data

SWE-smith

SWE-smith: Scaling Data for Software Engineering AgentsA toolkit and a 50k-instance dataset that turn Python repositories into executable software-engineering tasks for training coding agents, drawn from 128 GitHub repositories.

The loop

Language models do the data work: o3-mini injects bugs by modifying or rewriting functions and by mirroring reverted pull requests, an LM writes GitHub-style issue text for each failure, and SWE-agent running Claude 3.7 Sonnet solves the tasks to produce expert trajectories. 5,016 of those trajectories fine-tune Qwen 2.5 Coder 32B into SWE-agent-LM-32B, an open coding agent trained on tasks and solutions made by other agents.

The loop
SWE-smith
Dataset · swesmith.com
  1. o3-mini injects bugs and writes issues
  2. SWE-agent solves tasks for expert trajectories
  3. Trajectories fine-tune Qwen 2.5 Coder
  4. Open coding agent trained on agent data
  1. o3-mini injects bugs and writes issues
  2. SWE-agent solves tasks for expert trajectories
  3. Trajectories fine-tune Qwen 2.5 Coder
  4. Open coding agent trained on agent data
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Both task creation and expert solving are done by models, so a stronger agent trained on the data can take over generating and solving the next batch, removing the hundreds of hours of human curation the paper identifies as the bottleneck.

Evidence

SWE-agent-LM-32B reached 40.2% pass@1 on SWE-bench Verified, 33.4 points above its base model and state of the art among open-source models at release.

swe-agentssynthetic-taskscoding-agentsagent-trajectoriesdistillation