Disentangling Neural Identity from Encoder Identity in a Privacy Audit of BP-GPT’s fMRI-to-Text Decoding
Saketh Chebrolu ⋅ Rishabh Yadav ⋅ Joshua Seluzhitskiy ⋅ Nolan C Horvath ⋅ Kiran Nijjer
Abstract
Brain-to-text decoders map fMRI to language through learned embeddings that are smaller and easier to share than raw recordings, but whether these embeddings reveal whose brain produced them is unknown. We audit BP-GPT, an fMRI-to-text decoder, on eight participants retrained under a matched protocol. Raw functional connectivity identifies participants with 100\% top-1 accuracy (chance 12.5\%), and BP-GPT embeddings also reach 100\%, initially suggesting identity is fully preserved. However, BP-GPT trains a separate encoder per participant, so the model itself can leave a signature: cross-participant embeddings share essentially no stimulus information, and identification is already saturated at the encoder's first linear layer. Querying embeddings against profiles from independently seeded encoders isolates the neural signal: across twenty ordered pairs of five seed arms, accuracy falls to 54.3\% (subject-clustered 95\% CI 46.6-61.6\%; permutation $p=0.0001$), over four times chance but concentrated in a subset of participants. Within-model identification therefore overestimates re-identification risk in per-participant architectures. Privacy audits of released representations should distinguish fingerprints introduced by the data from those introduced by the model.
Chat is not available.
Successful Page Load