Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
A two-step training recipe is reported to make tamper localization hold up better across newer VLM image editors.
The paper targets pixel-level detection under shifts between systems such as ChatGPT, Gemini and Qwen-Image. Its method balances real and tampered samples inside each minibatch, then adds a small amount of support data from emerging VLM distributions late in training. The authors say this avoids training collapse and improves out-of-distribution robustness without overfitting to limited new domains. They report relative gains over PIXAR of 26.1% in average gIoU and 26.8% in cIoU across GPT-Images-2.0, Gemini-3.1, FLUX.2 and Seedream 4.5. ArXiv · AI/CL/LG's note
The paper targets pixel-level detection under shifts between systems such as ChatGPT, Gemini and Qwen-Image. Its method balances real and tampered samples inside each minibatch, then adds a small amount of support data from emerging VLM distributions late in training. The authors say this avoids training collapse and improves out-of-distribution robustness without overfitting to limited new domains. They report relative gains over PIXAR of 26.1% in average gIoU and 26.8% in cIoU across GPT-Images-2.0, Gemini-3.1, FLUX.2 and Seedream 4.5. ArXiv · AI/CL/LG's note
score 4