Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking
The paper tests a source-conditioned signal for spotting rewritten or synthesized reuse, not just AI authorship.
SCDG compares how a frozen language model predicts a suspicious document with and without a candidate source, turning the difference into token-level evidence of reuse. On PAN-derived benchmarks, the authors report 0.94 F1 for pairwise detection and stronger reranking results than their baselines. A same-topic Multi-News test flagged only 0.125% of pairs as source reuse under their calibrated setup, which the paper presents as evidence against confusing topical overlap with plagiarism. ArXiv · AI/CL/LG's note
SCDG compares how a frozen language model predicts a suspicious document with and without a candidate source, turning the difference into token-level evidence of reuse. On PAN-derived benchmarks, the authors report 0.94 F1 for pairwise detection and stronger reranking results than their baselines. A same-topic Multi-News test flagged only 0.125% of pairs as source reuse under their calibrated setup, which the paper presents as evidence against confusing topical overlap with plagiarism. ArXiv · AI/CL/LG's note
score 4