Alphabell.
Paper · Tools and harness

Self-optimizing Deep Research

Self-Optimizing Multi-Agent Systems for Deep ResearchA multi-agent deep research system (orchestrator, reader, aggregator and writer agents) whose prompts are optimized automatically by LLM-based optimizers instead of by hand.

The loop

LLM optimizers (TextGrad, GEPA and OpenAI's prompt optimizer) critique the deep research system's reports against expert rubrics and rewrite the prompt of one agent at a time. Each new system variant is scored by an LLM judge and better variants are kept, so the research system that runs the next queries is the one the loop produced.

The loop
Self-optimizing Deep Research
Paper · arXiv
  1. LLM optimizers critique deep research reports
  2. Optimizers rewrite research agent prompts
  3. Optimized system runs the next queries
  1. LLM optimizers critique deep research reports
  2. Optimizers rewrite research agent prompts
  3. Optimized system runs the next queries
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

A web research system tunes its own pipeline from judged outputs with no hand engineering, a step that can be repeated whenever the system or its optimizer improves.

Evidence

On ScholarQA computer science queries, starting from minimal prompts, GEPA with a custom meta-prompt reached a score of 0.705 versus 0.667 for expert-written prompts.

deep-researchprompt-optimizationmulti-agentllm-judgegepa