Decomposing Brain-LLM Alignment with Interpretable Semantic Representations
Camila Kolling ⋅ Mariya Toneva
Abstract
Large language models (LLMs) predict neural responses during naturalistic language comprehension, but the linguistic information underlying this brain alignment remains difficult to interpret. We use SCRIBE, an interpretable contextual semantic representation spanning $65$ neurocognitively grounded semantic properties, as a lens for decomposing brain--LLM alignment. Using variance partitioning between LLaMA2-7B and SCRIBE on naturalistic fMRI data, we separate brain-predictive variance shared between the two representations from variance unique to each. We find that shared variance is the dominant component across the whole brain and language-related regions, while both representations also retain significant unique contributions. We then examine whether the concreteness of individual semantic properties relates to shared and SCRIBE-unique brain-predictive variance. Concrete properties show greater shared variance with LLaMA2-7B than abstract properties across all examined regions, whereas abstract properties do not show greater SCRIBE-unique variance. Finally, property ablations reveal that semantic contributions extend across multiple dimensions, with a clearer cumulative effect for shared than SCRIBE-unique variance. Together, these results show that a substantial portion of brain–LLM alignment is associated with interpretable semantic information, while revealing complementary brain-predictive semantic information beyond the LLaMA2-7B representation.
Chat is not available.
Successful Page Load