Dialect Fingerprints in Language Model Internals: Label-Free Discovery and Controlled Dissection of Regional Features
Vasanth Sarathy ⋅ Timothy McKinnon
Abstract
Regional linguistic features often persist in a speaker's text regardless of topic or interlocutor: a Gulf Arabic speaker discussing Cairo typically retains regional markers in grammar and vocabulary, particularly those below the level of conscious awareness. We show that this meaning-independent fingerprint can be extracted from the internal representations of pretrained multilingual language models using techniques from mechanistic interpretability. Exploiting the parallel structure of the MADAR Arabic dialect corpus (25 city varieties expressing identical meanings), we subtract shared semantic content from model activations and train sparse autoencoders on the residuals without geographic supervision. The resulting features are geographically concentrated (278 significant at 4B scale), transfer to independently collected tweets, and respond causally to dialectal form. Controlled factorial dissection ($n = 20$ templates per cell) reveals that individual features encode dialectal variation at multiple levels of abstraction: one feature tracks the specific Egyptian WANT marker \emph{`aayiz} with near-perfect selectivity ($R^2 = 0.96$), while others respond to the volitional concept across four dialect families simultaneously ($p < 0.001$). These results establish that mechanistic interpretability can serve as a tool for computational dialectology, extracting structured, interpretable dialect fingerprints directly from model internals.
Chat is not available.
Successful Page Load