AdapTrack: Task Provenance from Merged LoRA Weights Alone
Mohammed Saad Shareef
Abstract
Merged LoRA adapters are published as weights, not as a list of the datasets behind them. We study whether that list is recoverable from the weights alone, as a question of dataset-level provenance rather than contributive attribution. Our adversary is passive: it observes only the merged parameter update — no queries to the model, no merge weights, no metadata, no cooperation from the publisher — and computes summary statistics of each module's singular spectrum. On a pool of 96 adapters spanning four task families, a classifier over those statistics determines whether a given family was among a merge's constituents at AUC $0.969$ and $0.975$ for two targets, with true-positive rates of $0.85$ and $0.90$ at zero false positives, under adapter-level holdout, cluster bootstrap intervals over adapters, and a permutation null. Examining what merging preserves, we find that for linear merges only total per-module magnitude passes through composition (median within-module $R^2 = 0.96$); every finer spectral statistic falls below $0.41$. We then evaluate two structurally different defenses against an adaptive attacker who retrains on defended features; neither reduces the attack below $0.87$ AUC, against an undefended baseline of $0.93$ on the same subset. Evaluated against a fixed attacker instead, one of them appears to drive the attack below chance, overstating protection by $0.53$ AUC.
Chat is not available.
Successful Page Load