Reading Documents out of Weight Updates: Hypernetwork-Written LoRA Adapters Linearly Encode Their Document
Aryan Keluskar
Abstract
Hypernetworks such as Doc-to-LoRA write an entire document into a low-rank weight update in a single forward pass. We ask whether the document can be read back out, and find that it can, but only when a hypernetwork wrote it. A linear probe from the low-rank factors of a Doc-to-LoRA update to the document's bag-of-words identifies the held-out document at recall@5 $=0.86$--$0.99$ across three model families (chance $<0.16$), whereas ordinary SGD adapters that internalise the same documents are decodable only at chance, a $10$--$20\sigma$ gap at matched effective update norm. A Doc-to-LoRA update is a near-linear image of the document embedding (held-out $R^2{=}0.97$) confined to a low-dimensional subspace; an SGD update is not a function of any shared embedding. The leak does not even require the weights: an attacker who can only \emph{query} a deployed adapter recovers its document from output logits alone (recall@5 $0.94$, and $0.50$ when the API exposes only its top-$20$ logprobs), and this behavioural leak is \emph{universal}: even an SGD adapter, unreadable from its weights, regurgitates its document when queried ($0.59$). The readability of a weight update thus depends on its generator, and any deployed document fine-tune leaks its training text to anyone who can query it.
Chat is not available.
Successful Page Load