Megadose AI progress, ranked and analyzed.

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

· ArXiv · AI/CL/LG ·
The models often held usable harmful-meme signals internally, but failed to route them into the final classification.

The paper tests Gemma-3 and Qwen3.5 across six harmful-content benchmarks, plus Spanish and Hindi-English code-mixed evaluations. Sparse readouts beat the models’ native predictions on all six primary binary tasks, with Qwen showing the largest gap. The authors argue the gap reflects supervised accessibility rather than an already-formed native decision rule. Calibration-only routing recovered 93.3% of the mean gap, while probe-distilled LoRA improved native predictions but multi-task adaptation caused negative transfer. ArXiv · AI/CL/LG's note

score 5

Categories: Research