Multimodal Model Diffing for Feature Discovery and Control
MMDiff turns sparse autoencoders into controls for specific multimodal model behaviors.
The paper says the method compares base language-model features with multimodal-adapted ones to isolate what visual training changed. It then uses token-level contrastive analysis to find causal features for tasks including spatial reasoning, safety attacks, and OCR. Removing those features selectively weakens the targeted behavior, while steering them improves spatial and OCR accuracy over a single-layer steering baseline. ArXiv · AI/CL/LG's note
The paper says the method compares base language-model features with multimodal-adapted ones to isolate what visual training changed. It then uses token-level contrastive analysis to find causal features for tasks including spatial reasoning, safety attacks, and OCR. Removing those features selectively weakens the targeted behavior, while steering them improves spatial and OCR accuracy over a single-layer steering baseline. ArXiv · AI/CL/LG's note
score 5