Pretraining Data Can Be Poisoned through Computational Propaganda
The paper identifies public discussion interfaces as a plausible route for poisoning language-model pretraining data.
The authors argue that attacks need not rely on editing major sources like Wikipedia. They introduce HalfLife, an analysis method for estimating whether adversarial web content survives crawling and curation into training corpora. Their abstract frames third-party webpage content as a possible attack vector for pretraining pipelines. ArXiv · AI/CL/LG's note
The authors argue that attacks need not rely on editing major sources like Wikipedia. They introduce HalfLife, an analysis method for estimating whether adversarial web content survives crawling and curation into training corpora. Their abstract frames third-party webpage content as a possible attack vector for pretraining pipelines. ArXiv · AI/CL/LG's note
score 6