PMDA: Mechanistic Data Attribution for Steerable Pluralism
Abstract
Pluralistic alignment aims to accommodate heterogeneous values, yet little is known about which pretraining data support pluralistic behavior. Mechanistic data attribution (MDA) offers a route from model behavior to influential training sequences, but its rankings require causal validation. We introduce Pluralistic Mechanistic Data Attribution (PMDA), a prospective protocol for mechanism-guided data selection. The proposed study first validates a Moral Integrity Corpus (MIC) behavioral measure for a Pythia base model, then, conditional on passing that gate, identifies and locks an attention-head mechanism before attributing it to Pile sequences. Three proposed criteria assess ranking stability, utility under matched replay, and bounded general-capability cost. An exploratory feasibility appendix evaluates fixed Pythia and OLMo base checkpoints without parameter updates or internal interventions; its results do not establish the behavioral gate. Mechanism discovery, influence ranking, and replay remain prospective.