Research and evals
Research automation
AI that proposes hypotheses, runs experiments and writes up machine learning research.
autoresearch
- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat
Research automationAndrej Karpathy · 2026
autoresearch-distillation
- Qwen3-14B edits GPT training script
- Edit is trained and scored
- Score rewards Qwen3-14B weights update
- Trained checkpoint re-run in autoresearch loop
- repeat
Research automationExperiential Labs (Naihin, Fallah) · 2026
DeepScientist
- LLM agents propose hypotheses on AI tasks
- Agents implement and test them
- Findings Memory steers later proposals
- Validated finding directly improves AI system
- repeat
Research automationWeng et al. (Westlake University) · 2025
The AI Scientist-v2
- LLM agents propose ML research ideas
- Agents run experiments with tree search
- Agents write up results and papers
- Output is new knowledge about models
- repeat
Research automationYamada et al. (Sakana AI) · 2025