Self-optimizing Deep Research
Self-Optimizing Multi-Agent Systems for Deep ResearchA multi-agent deep research system (orchestrator, reader, aggregator and writer agents) whose prompts are optimized automatically by LLM-based optimizers instead of by hand.
LLM optimizers (TextGrad, GEPA and OpenAI's prompt optimizer) critique the deep research system's reports against expert rubrics and rewrite the prompt of one agent at a time. Each new system variant is scored by an LLM judge and better variants are kept, so the research system that runs the next queries is the one the loop produced.
- LLM optimizers critique deep research reports
- Optimizers rewrite research agent prompts
- Optimized system runs the next queries
- LLM optimizers critique deep research reports
- Optimizers rewrite research agent prompts
- Optimized system runs the next queries
Why it is a road to recursion
A web research system tunes its own pipeline from judged outputs with no hand engineering, a step that can be repeated whenever the system or its optimizer improves.
Evidence
On ScholarQA computer science queries, starting from minimal prompts, GEPA with a custom meta-prompt reached a score of 0.705 versus 0.667 for expert-written prompts.
Related loops
More tools and harness →- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat
- Reflection LLM writes revised module prompts
- Winning prompt candidates are kept
- Optimized prompts return to same system
- New system traces feed next reflection round
- repeat
- Meta agent programs new agent designs
- Designs are evaluated on target tasks
- Results are added to an archive
- Meta agent uses archive for next design
- repeat