When Specifications Drift: Measuring Semantic Preservation Across LLM Representations
Abstract
Verification establishes that an implementation satisfies a specification; it does not establish that a model-translated specification preserves the intended meaning. We study this gap in a controlled equational domain. Seventeen deterministic encodings and a canonical control present the same 500 magma identities to nine language models. A formal implication oracle distinguishes equivalence, strict weakening, strict strengthening, and incomparability when outputs are within its certified catalogue; coverage failures are reported separately. Across 153,000 recorded calls, including repeats and three follow-up studies, certified recovery ranges from 73.7% for mechanical bracket-word text to 8.5% for postfix notation. On confusable identifiers, 30.3% of main responses are proven strictly weaker, and 93.7% of those use fewer distinct variables. A six-scheme identifier study reduces recovery from 64.4% to 3.1% at Latin/Cyrillic homoglyphs. Reading instructions with examples can recover large losses, while an anti-transformation instruction substantially improves even a verbatim-input control. Repeated sampling confirms substantial residual variability. These results motivate evaluating specification translation by directional semantic relations and explicit oracle coverage, rather than syntax or a single accuracy score.