Information Parity for Code: The Scope of Transfer in Multilingual Code Models
Alexander Tsvetkov ⋅ Alon Kipnis
Abstract
LLMs exhibit uneven performance across programming languages, yet benchmark gaps often reflect narrow task regimes and evaluation-specific confounds. This makes it hard to tell whether poor performance reflects a broader weakness with a given language, a broader difficulty on a given task, or a limitation specific to that task-language combination. We adapt Information Parity, a measure of relative cross-language representational efficiency developed for natural language, to code. We define Code Information Parity (Code IP) on a controlled parallel corpus of 22 algorithms across 20 languages, audited to preserve core semantics while minimizing boilerplate and excluding library shortcuts. The results are regime-specific: Code IP correlates strongly with multilingual benchmark performance on short-horizon generation ($\rho$ from $-0.55$ to $-0.81$), remains negative though moderated on execution prediction ($\rho=-0.429$), and disappears on whole-program synthesis ($\rho=+0.105$). Rather than a universal metric of multilingual coding ability, Code IP isolates a reliable cross-language signal on short-horizon generation and delineates where that signal no longer separates cleanly from broader task and evaluation effects.
Chat is not available.
Successful Page Load