Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
The paper targets authorship at the token level, not just whether a whole document looks AI-written.
The authors propose smoothing token-level detection scores so likely LLM-generated spans can be localized inside mixed human-LLM text. The method uses an adaptive Lepski-type rule to choose bandwidth based on local authorship structure. They say it does not require token-level labeled training data, and report strong results on synthetic and realistic datasets. They also note a public website implementing the method. ArXiv · AI/CL/LG's note
The authors propose smoothing token-level detection scores so likely LLM-generated spans can be localized inside mixed human-LLM text. The method uses an adaptive Lepski-type rule to choose bandwidth based on local authorship structure. They say it does not require token-level labeled training data, and report strong results on synthetic and realistic datasets. They also note a public website implementing the method. ArXiv · AI/CL/LG's note
score 4