AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
The loop: The AREX-2 agent synthesizes long-horizon improvement trajectories from ML engineering tasks and uses them to train its successor for better performance on MLE-bench.
AREX-2 trains an agent on synthesized long-horizon improvement trajectories from ML and programming tasks to advance its self-improving capabilities.







