Alphabell.
Paper · Tools and harness

ScholarEvolve

Learning from Research: Toward Lifelong Agent Harness Evolution (ScholarEvolve)A framework in which an LLM research agent retrieves and reads AI papers and turns their mechanisms into changes to an agent's harness (tools, context management, skills, memory, workflows), which a coding agent then implements and tests.

The loop

A research agent audits the task agent's failures, turns them into capability gaps, retrieves papers on those gaps and writes mutation blueprints, and a coding agent implements each one as a harness module change while the model weights stay fixed. Candidates and their combinations are validated against the current champion harness, the winner becomes the new champion, and new publications can start further generations.

The loop
ScholarEvolve
Paper · arXiv
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

AI agents read the latest AI research to rebuild an AI agent's harness, so newly published methods reach the agent without a human engineer translating them, and the process repeats as the literature grows.

Evidence

The evolved harness raised Qwen3.5-27B task goal completion on AppWorld Challenge from 49.6% to 63.6% and GPT-5.4-mini pass@1 on Tau2-Bench Telecom from 72.7% to 81.9%.

agent-harnessliterature-miningresearch-agentevolutionlifelong-learning