PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs
Luze Sun ⋅ Anshuman Suri ⋅ Harsh Chaudhari ⋅ Cristina Nita-Rotaru ⋅ Alina Oprea
Abstract
When practitioners fine-tune LLMs on unvetted datasets, an adversary can exploit the data supply chain through task-level poisoning: inserting a small number of crafted instruction-response pairs that cause the model to embed attacker-specified entities, such as a country, in outputs for a target task while limiting their occurrence on other tasks. We introduce PoisonForge, a benchmark that parameterizes this threat along four dimensions (bias type, poisoning mode, appearance count, and target length) and evaluates 12 open-weight models (from 2B to 32B parameters) across five families under a default 1\% poison budget. With only 10 poisoned examples added to 1{,}000 benign examples, the mean attack success rate (ASR) across 12 models and 16 configurations is $38.8\%$, with 11 of 12 exceeding $70\%$ ASR in at least one configuration. We also measure unintended spillover to non-target tasks and performance on standard capability benchmarks. We observe that five appearances yield higher ASR than one in the main configurations, poisoning-mode differences depend on bias type, and longer target lengths generally yield lower ASR. Within a target task, configuration-level patterns are broadly shared across models, while absolute ASR does not transfer reliably across target tasks and configuration rankings are partly shared. We release all configurations, pipelines, and analysis code to support reproducible comparisons.
Chat is not available.
Successful Page Load