Megadose AI progress, ranked and analyzed.

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

· ArXiv · AI/CL/LG ·
The paper targets authorship at the token level, not just whether a whole document looks AI-written.

The authors propose smoothing token-level detection scores so likely LLM-generated spans can be localized inside mixed human-LLM text. The method uses an adaptive Lepski-type rule to choose bandwidth based on local authorship structure. They say it does not require token-level labeled training data, and report strong results on synthetic and realistic datasets. They also note a public website implementing the method. ArXiv · AI/CL/LG's note

score 4

Categories: Research