Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents
The paper claims agent watermarks need to survive rewritten tool/action labels and forged trajectories, not just token changes.
SBW moves the watermark from exact action symbols to semantic action clusters conditioned on history, then uses keyed collision-resistant binning to make fresh bucket choices unpredictable. In the reported tests, cluster-level detection stayed much higher than exact-symbol methods under rewriting on ToolBench and ALFWorld, while preserving stronger action agreement than logit biasing. The authors say keyed binning cut adaptive forgery from 100% to the false-positive floor at their main setting. They also flag an unresolved limit: chained replay of a victim’s own steps still verified at high rates across models. ArXiv · AI/CL/LG's note
SBW moves the watermark from exact action symbols to semantic action clusters conditioned on history, then uses keyed collision-resistant binning to make fresh bucket choices unpredictable. In the reported tests, cluster-level detection stayed much higher than exact-symbol methods under rewriting on ToolBench and ALFWorld, while preserving stronger action agreement than logit biasing. The authors say keyed binning cut adaptive forgery from 100% to the false-positive floor at their main setting. They also flag an unresolved limit: chained replay of a victim’s own steps still verified at high rates across models. ArXiv · AI/CL/LG's note
score 5