Pretraining Data Can Be Poisoned through Computational Propaganda
Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally, to measure
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Hackers Abuse SEO Poisoning and Hidden HTML to Trick AI Agents Into Following Malicious Instructions - CyberSecurityNews →
- PossiblePossibly related (embedding) · 49%Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise →
- PossiblePossibly related (embedding) · 48%Reddit deploys AI to cut spam and harmful content exposure - MSN →
- PossiblePossibly related (embedding) · 48%Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign) →
- PossiblePossibly related (embedding) · 48%Reddit Deploys LLMs to Fight AI-Generated Spam, Cutting Exposure by 20% - finance.biggo.com →
- FuzzyOverlapping authors or contributors · 62%deepspeedai/DeepSpeed →
“Shared author/contributor keys: smith”
- LinkedLinked via arxiv author · 85%Victoria Graf →
“Pretraining Data Can Be Poisoned through Computational Propaganda”
- LinkedLinked via arxiv author · 85%Hannaneh Hajishirzi →
“Pretraining Data Can Be Poisoned through Computational Propaganda”
