DeepScientist
DeepScientist: Advancing Frontier-Pushing Scientific Findings ProgressivelyDeepScientist is a goal-directed autonomous research system that runs month-long discovery campaigns on frontier AI tasks, framing discovery as Bayesian optimization over a hierarchy of hypothesize, verify and analyze steps.
LLM agents propose hypotheses for beating the state of the art on human-chosen AI tasks (agent failure attribution, LLM inference acceleration, AI-generated text detection), implement and test them, and record outcomes in a cumulative Findings Memory that steers later proposals, promoting promising findings to higher-fidelity validation. Because the tasks are AI methods, a validated finding is a direct improvement to an AI system, such as faster LLM inference.
- LLM agents propose hypotheses on AI tasks
- Agents implement and test them
- Findings Memory steers later proposals
- Validated finding directly improves AI system
- LLM agents propose hypotheses on AI tasks
- Agents implement and test them
- Findings Memory steers later proposals
- Validated finding directly improves AI system
Why it is a road to recursion
It shows an AI system producing measurable improvements to AI methods over long horizons, including LLM inference speed, which is the work a recursive loop would compound.
Evidence
Using over 20,000 GPU hours, the system generated about 5,000 ideas, validated about 1,100 experimentally, and its abstract reports surpassing human-designed SOTA methods on three frontier AI tasks by 183.7%, 1.9% and 7.9%.
Related loops
More research automation →- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat
- Qwen3-14B edits GPT training script
- Edit is trained and scored
- Score rewards Qwen3-14B weights update
- Trained checkpoint re-run in autoresearch loop
- repeat
- LLM agents propose ML research ideas
- Agents run experiments with tree search
- Agents write up results and papers
- Output is new knowledge about models
- repeat