Multimodal Model Diffing for Feature Discovery and Control
MMDiff turns multimodal sparse autoencoders into controls, not just inspection tools.
The paper says the method compares a base language model SAE with its multimodal-adapted version to find features changed by visual training. It then uses contrastive firing analysis to isolate causal task features and remove or steer them. Across LLaVA-MORE, PaliGemma 2, and InternVL3.5, removing discovered features selectively hurt spatial and OCR behavior and cut multimodal safety attack success without affecting VQA. Steering those features beat a single-layer steering baseline on spatial and OCR accuracy. HF Daily Papers' note
The paper says the method compares a base language model SAE with its multimodal-adapted version to find features changed by visual training. It then uses contrastive firing analysis to isolate causal task features and remove or steer them. Across LLaVA-MORE, PaliGemma 2, and InternVL3.5, removing discovered features selectively hurt spatial and OCR behavior and cut multimodal safety attack success without affecting VQA. Steering those features beat a single-layer steering baseline on spatial and OCR accuracy. HF Daily Papers' note
score 5