Alphabell.
Paper · Interpretability and oversight

MAIA

A Multimodal Automated Interpretability AgentMAIA gives a pretrained vision-language model tools to run interpretability experiments on other vision models (synthesizing and editing images, finding maximally activating exemplars, summarizing results) to describe features and find failure modes.

The loop

A vision-language model agent designs and runs experiments on a ResNet's final-layer neurons and decides which ones rely on spurious background cues. Its selection is then used to retrain the classifier's final layer, so the interpreting model directly improves the robustness of the model it interprets.

The loop
MAIA
Paper · arXiv
  1. Agent runs interpretability experiments
  2. Agent finds spurious background cues
  3. Selection retrains classifier final layer
  4. Improves robustness of interpreted model
  1. Agent runs interpretability experiments
  2. Agent finds spurious background cues
  3. Selection retrains classifier final layer
  4. Improves robustness of interpreted model
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It runs an interpret-then-fix cycle with an AI agent doing the interpretation, which can be pointed at larger models as the agent's backbone improves.

Evidence

On Spawrious, a final layer retrained on 22 MAIA-selected neurons reached 0.837 balanced test accuracy versus 0.731 using all 512 units and 0.757 for 22 l1-selected neurons chosen without balanced data.

interpretabilityagentvision-modelsspurious-featurestool-use