Tongyi DeepResearch
An open 30.5B-parameter (3.3B active) web research agent trained end to end with agentic continual pretraining, supervised finetuning and reinforcement learning on data from a fully automatic synthesis pipeline.
A fully automated, model-driven data engine writes the agent's training data: AgentFounder synthesizes agentic pretraining data and feeds data from the post-training pipeline back in as a data flywheel, and a question-crafting agent with search, retrieval and Python tools repeatedly upgrades seed questions into harder research tasks. During RL, a synthesis and filtering pipeline adjusts the training set in real time from training dynamics, which the team describes as closing the loop between data generation and model training.
- Agent synthesizes agentic pretraining data
- Pipeline adjusts training set in real time
- Data feeds back into model training
- Agent synthesizes agentic pretraining data
- Pipeline adjusts training set in real time
- Data feeds back into model training
Why it is a road to recursion
The training data, the task curriculum and the RL environment (a simulated offline-Wikipedia setup) are all produced by models and tuned to the trainee's progress, so agent capability and data difficulty can rise together without human annotation.
Evidence
The model scores 32.9 on Humanity's Last Exam, 43.4 on BrowseComp and 46.7 on BrowseComp-ZH.
Related loops
More training data →- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat
- Language model generates a self-edit
- Edit is applied as weight update
- RL rewards self-edit by downstream performance
- repeat
- o3-mini injects bugs and writes issues
- SWE-agent solves tasks for expert trajectories
- Trajectories fine-tune Qwen 2.5 Coder
- Open coding agent trained on agent data
- repeat