Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
The paper says AI-written research can be spotted by broken scientific reasoning, not just by telltale wording.
The authors introduce SciSlopBench, pairing 390 AI-generated papers with human-written papers matched by research problem and contribution type. Their six measures identify the AI paper in each pair with 85.9% accuracy, outperforming Binoculars in the abstract. They also report that higher “scientific slop” tracks with lower ICLR ratings and separates rejected from accepted papers above chance from 2017 to 2025. Their proposed SciSlopHarness reduces the remaining AI-human gap by 63% over the strongest revision baseline by limiting revisions to changes supported by experiment records. HF Daily Papers' note
The authors introduce SciSlopBench, pairing 390 AI-generated papers with human-written papers matched by research problem and contribution type. Their six measures identify the AI paper in each pair with 85.9% accuracy, outperforming Binoculars in the abstract. They also report that higher “scientific slop” tracks with lower ICLR ratings and separates rejected from accepted papers above chance from 2017 to 2025. Their proposed SciSlopHarness reduces the remaining AI-human gap by 63% over the strongest revision baseline by limiting revisions to changes supported by experiment records. HF Daily Papers' note
score 5