Alphabell.
Paper · Research automation

DeepScientist

DeepScientist: Advancing Frontier-Pushing Scientific Findings ProgressivelyDeepScientist is a goal-directed autonomous research system that runs month-long discovery campaigns on frontier AI tasks, framing discovery as Bayesian optimization over a hierarchy of hypothesize, verify and analyze steps.

The loop

LLM agents propose hypotheses for beating the state of the art on human-chosen AI tasks (agent failure attribution, LLM inference acceleration, AI-generated text detection), implement and test them, and record outcomes in a cumulative Findings Memory that steers later proposals, promoting promising findings to higher-fidelity validation. Because the tasks are AI methods, a validated finding is a direct improvement to an AI system, such as faster LLM inference.

The loop
DeepScientist
Paper · arXiv
  1. LLM agents propose hypotheses on AI tasks
  2. Agents implement and test them
  3. Findings Memory steers later proposals
  4. Validated finding directly improves AI system
  1. LLM agents propose hypotheses on AI tasks
  2. Agents implement and test them
  3. Findings Memory steers later proposals
  4. Validated finding directly improves AI system
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It shows an AI system producing measurable improvements to AI methods over long horizons, including LLM inference speed, which is the work a recursive loop would compound.

Evidence

Using over 20,000 GPU hours, the system generated about 5,000 ideas, validated about 1,100 experimentally, and its abstract reports surpassing human-designed SOTA methods on three frontier AI tasks by 183.7%, 1.9% and 7.9%.

automated-researchbayesian-optimizationllm-agentsinference-accelerationfrontier-tasks