Megadose AI progress, ranked and analyzed.

Multimodal Model Diffing for Feature Discovery and Control

· HF Daily Papers ·
MMDiff turns multimodal sparse autoencoders into controls, not just inspection tools.

The paper says the method compares a base language model SAE with its multimodal-adapted version to find features changed by visual training. It then uses contrastive firing analysis to isolate causal task features and remove or steer them. Across LLaVA-MORE, PaliGemma 2, and InternVL3.5, removing discovered features selectively hurt spatial and OCR behavior and cut multimodal safety attack success without affecting VQA. Steering those features beat a single-layer steering baseline on spatial and OCR accuracy. HF Daily Papers' note

score 5

Categories: Research