Invisible Shortcuts: Why Vision Encoders Know Your Camera
Vision encoders can learn camera and processing traces that humans do not see.
The paper says large-scale pretraining can turn pixel-level metadata artifacts into predictive features when those traces correlate with labels or captions. In controlled tests, stronger metadata-semantics links made models more sensitive to those traces and more brittle when metadata distributions shifted. The authors also report mitigation methods that reduce sensitivity to both targeted and unseen metadata without hurting downstream performance. That same sensitivity may help explain why some encoders are good at detecting generated images. HF Daily Papers' note
The paper says large-scale pretraining can turn pixel-level metadata artifacts into predictive features when those traces correlate with labels or captions. In controlled tests, stronger metadata-semantics links made models more sensitive to those traces and more brittle when metadata distributions shifted. The authors also report mitigation methods that reduce sensitivity to both targeted and unseen metadata without hurting downstream performance. That same sensitivity may help explain why some encoders are good at detecting generated images. HF Daily Papers' note
score 5