Alphabell.
Project · Training data

Tongyi DeepResearch

An open 30.5B-parameter (3.3B active) web research agent trained end to end with agentic continual pretraining, supervised finetuning and reinforcement learning on data from a fully automatic synthesis pipeline.

The loop

A fully automated, model-driven data engine writes the agent's training data: AgentFounder synthesizes agentic pretraining data and feeds data from the post-training pipeline back in as a data flywheel, and a question-crafting agent with search, retrieval and Python tools repeatedly upgrades seed questions into harder research tasks. During RL, a synthesis and filtering pipeline adjusts the training set in real time from training dynamics, which the team describes as closing the loop between data generation and model training.

The loop
Tongyi DeepResearch
Project · GitHub
  1. Agent synthesizes agentic pretraining data
  2. Pipeline adjusts training set in real time
  3. Data feeds back into model training
  1. Agent synthesizes agentic pretraining data
  2. Pipeline adjusts training set in real time
  3. Data feeds back into model training
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

The training data, the task curriculum and the RL environment (a simulated offline-Wikipedia setup) are all produced by models and tuned to the trainee's progress, so agent capability and data difficulty can rise together without human annotation.

Evidence

The model scores 32.9 on Humanity's Last Exam, 43.4 on BrowseComp and 46.7 on BrowseComp-ZH.

research-agentssynthetic-dataagentic-rldata-flywheelopen-weights