Towards a Physical Theory of Cross-Lingual Representation Alignment: The RG Hypothesis
Santiago Acevedo ⋅ Giulio Biroli ⋅ Alessandro Laio ⋅ Marco Baroni
Abstract
It was recently observed that language representations of LLMs processing sets of translated sentences develop similar neighborhood structures, namely, they become locally aligned. However, we still lack an understanding of what information is shared by representations in semantic correspondence. We propose the Renormalization Group from Statistical Mechanics as a framework for modeling the information shared between translations. Consistent with this picture, we argue that natural text and LLM neural activity show signatures of scale invariance, and we show that cross-lingual representational convergence is strongly concentrated in the long-wavelength modes of LLM neural activity along the text axis.
Chat is not available.
Successful Page Load