SWE-smith
SWE-smith: Scaling Data for Software Engineering AgentsA toolkit and a 50k-instance dataset that turn Python repositories into executable software-engineering tasks for training coding agents, drawn from 128 GitHub repositories.
Language models do the data work: o3-mini injects bugs by modifying or rewriting functions and by mirroring reverted pull requests, an LM writes GitHub-style issue text for each failure, and SWE-agent running Claude 3.7 Sonnet solves the tasks to produce expert trajectories. 5,016 of those trajectories fine-tune Qwen 2.5 Coder 32B into SWE-agent-LM-32B, an open coding agent trained on tasks and solutions made by other agents.
- o3-mini injects bugs and writes issues
- SWE-agent solves tasks for expert trajectories
- Trajectories fine-tune Qwen 2.5 Coder
- Open coding agent trained on agent data
- o3-mini injects bugs and writes issues
- SWE-agent solves tasks for expert trajectories
- Trajectories fine-tune Qwen 2.5 Coder
- Open coding agent trained on agent data
Why it is a road to recursion
Both task creation and expert solving are done by models, so a stronger agent trained on the data can take over generating and solving the next batch, removing the hundreds of hours of human curation the paper identifies as the bottleneck.
Evidence
SWE-agent-LM-32B reached 40.2% pass@1 on SWE-bench Verified, 33.4 points above its base model and state of the art among open-source models at release.
Related loops
More training data →- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat
- Agent synthesizes agentic pretraining data
- Pipeline adjusts training set in real time
- Data feeds back into model training
- repeat
- Language model generates a self-edit
- Edit is applied as weight update
- RL rewards self-edit by downstream performance
- repeat