Alphabell.
Project · Tools and harness

GEPA

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement LearningA prompt optimizer in which an LLM reads execution traces of an AI system, diagnoses failures in natural language and proposes prompt updates, keeping a Pareto frontier of candidates. Released as a library and integrated into DSPy.

The loop

A reflection LLM reads trajectories (reasoning, tool calls, tool outputs) from an LLM system, writes revised prompts for its modules and tests them, keeping candidates that win on some examples on a Pareto frontier and combining their lessons. The optimized prompts go back into the same system, whose new traces feed the next round of reflection.

The loop
GEPA
Project · GitHub
  1. Reflection LLM writes revised module prompts
  2. Winning prompt candidates are kept
  3. Optimized prompts return to same system
  4. New system traces feed next reflection round
  1. Reflection LLM writes revised module prompts
  2. Winning prompt candidates are kept
  3. Optimized prompts return to same system
  4. New system traces feed next reflection round
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It replaces weight updates with language-level self-diagnosis that an LLM system can apply to its own prompts and code, a cheap step that can be rerun each time the underlying model improves.

Evidence

Across six tasks GEPA beat GRPO by 6% on average and by up to 20% while using up to 35x fewer rollouts, and beat the MIPROv2 prompt optimizer by over 10%.

prompt-optimizationreflectiondspyparetorl-alternative