MAIA
A Multimodal Automated Interpretability AgentMAIA gives a pretrained vision-language model tools to run interpretability experiments on other vision models (synthesizing and editing images, finding maximally activating exemplars, summarizing results) to describe features and find failure modes.
A vision-language model agent designs and runs experiments on a ResNet's final-layer neurons and decides which ones rely on spurious background cues. Its selection is then used to retrain the classifier's final layer, so the interpreting model directly improves the robustness of the model it interprets.
- Agent runs interpretability experiments
- Agent finds spurious background cues
- Selection retrains classifier final layer
- Improves robustness of interpreted model
- Agent runs interpretability experiments
- Agent finds spurious background cues
- Selection retrains classifier final layer
- Improves robustness of interpreted model
Why it is a road to recursion
It runs an interpret-then-fix cycle with an AI agent doing the interpretation, which can be pointed at larger models as the agent's backbone improves.
Evidence
On Spawrious, a final layer retrained on 22 MAIA-selected neurons reached 0.837 balanced test accuracy versus 0.731 using all 512 units and 0.757 for 22 l1-selected neurons chosen without balanced data.
Related loops
More interpretability and oversight →- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
- Auditor model tests target model
- LLM judge scores target behavior
- Feeds into next Claude model release
- repeat
- CriticGPT critiques ChatGPT code answers
- Humans catch errors using critiques
- Improves labels for RLHF training
- Labels train the next ChatGPT
- repeat