ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
AI-text detectors that look strong on direct LLM output failed much more often when the input was human text rewritten by an LLM.
The paper introduces ARB, a benchmark built from 1,800 human source texts and four open-weight generators. Each item is matched across human text, direct LLM generation, LLM-rewritten human text, and LLM-rewritten LLM text. At a strict 1% false-positive setting, FastDetectGPT and Binoculars-falcon-7b caught more than 90% of direct LLM text but only 30.8% and 15.1% of LLM-rewritten human text. The authors say conventional human-vs-LLM benchmarks do not predict detector behavior on human-authored work revised by an LLM. ArXiv · AI/CL/LG's note
The paper introduces ARB, a benchmark built from 1,800 human source texts and four open-weight generators. Each item is matched across human text, direct LLM generation, LLM-rewritten human text, and LLM-rewritten LLM text. At a strict 1% false-positive setting, FastDetectGPT and Binoculars-falcon-7b caught more than 90% of direct LLM text but only 30.8% and 15.1% of LLM-rewritten human text. The authors say conventional human-vs-LLM benchmarks do not predict detector behavior on human-authored work revised by an LLM. ArXiv · AI/CL/LG's note
score 4