Cognitive Firewalls: A Synthetic Account of Cross-Lingual Reasoning Collapse
Abstract
The reasoning capabilities of large language models are reported to drop sharply when they are forced to think in heavily regulated target languages such as Chinese — a phenomenon the literature attributes to data imbalance, multilingual transfer failures, or, more recently, RL-induced cross-lingual collapse. We propose and test an alternative mechanistic account: that the gap reflects RL-induced selection between conditional policies sharing the same parameters, rather than a loss of underlying capability. We construct a model organism for this account. Two policies — identified solely by the language of an unsupervised chain of thought — are installed via SFT on disjoint splits of ARC-Challenge, with one policy trained to answer and the other to refuse on each split. Reinforcement learning with a final-answer reward, restricted to one split, drives the language of reasoning to collapse onto the rewarded policy across both splits, producing a 67% → 10% accuracy collapse on the held-out split despite no reward signal ever observing it. General capability is unaffected: MMLU accuracy is 71.2% before and after RL. A depth-wise activation steering intervention then recovers held-out accuracy from 0% to 85% on the worst-case slice (Chinese-CoT completions) while preserving Chinese as the surface chain-of-thought language and leaving the rewarded policy intact (83% vs. 94% unsteered). The capability was never lost; it was gated. To the extent deployed regulated-language models contain analogous structure, surface evaluations of such models will systematically understate capability hidden behind learned policy gates.