Same Score, Different Evidence: Decodability, Surface Sufficiency, and Causal Relevance in Code Models
Naing Oo Lwin ⋅ Ark Dutt ⋅ Sreedhyuti Nimmagadda ⋅ Michael Ji ⋅ Randy Lim ⋅ Cole Blondin ⋅ Archana Vaidheeswaran
Abstract
Interpretability experiments must distinguish model mechanisms from measurement artifacts. We present a case-study-backed reporting matrix for linear probes over five identifier properties and roles in code models. The study includes three models and collectively covers all seven XLCoST programming languages without forming a complete factorial design. The matrix binds each claim to an estimand, comparator, matching unit, uncertainty statement, outcome, and falsifier. Two matched cases motivate it. On Python, boolean occurrence type is decodable at $0.981$–$0.988$ macro-F1, while a masked source-line classifier reaches $0.983$ without a language model. Paired problem-clustered deltas are $-0.003$ to $+0.002$, with intervals excluding probe advantages above $0.014$–$0.017$, although the sign of that difference varies by language and by comparator. A full-residual patch at a class or structure query-name site also fails its specified bidirectional gate in Qwen2.5-1.5B, recovering $0.009$ to $0.020$ of the matched behavioral gap. This bounds the effect at that site rather than showing global non-use. A third case exposes a failed control. Role-conditioned iterator renaming yields $0.998$–$0.999$ F1 at embedding layer zero, consistent with a lexical code introduced by the intervention rather than semantic robustness. Decodability, capacity control, cue sensitivity, surface sufficiency, and site-state causal relevance therefore require separate, outcome-aware records. We make the implementation of our methods accessible at anonymous.4open.science/r/SameScoreDifferentEvidence2026.
Chat is not available.
Successful Page Load