Validity Before Verdicts: Measuring Hidden Communication in Language Models
Abstract
An output from a language model can carry decodable bits and pass a monitor (a detector for hidden messages) yet fail its assigned task or deliver only part of a hidden message. A positive measured rate alone therefore does not establish usable hidden communication. We present Supported Operational Rate (SOR) v2.1, a protocol for deciding when an experiment supports an operational claim about a covert channel. It checks execution integrity, task competence, measurement validity, and full-message recovery. It reports SUPPORTED when requirements pass, REJECTED for operational failure under valid measurement, INCONCLUSIVE for unresolved evidence, and INVALID for broken execution; each has a distinct rate-reporting rule. We apply SOR in a prospectively frozen experiment with a pinned Qwen2.5-1.5B sender, 200 factual-QA tasks, and an arithmetic channel at two payload sizes. A preregistered comparator that reads the nominal rate without the full validity checks reports 0.0286 bits per emitted token (95% CI [0.0218, 0.0358]) at 2 bits and 0.0443 ([0.0331, 0.0566]) at 8 bits. On the same outputs, SOR returns REJECTED for both cells because honest task accuracy and exact message recovery fall below the frozen requirements; encoded task accuracy also fails at 2 bits. The monitor passes its preregistered point-estimate checks, but its held-out false-positive rate of 0.084 has a 95% interval of [0.040, 0.136] crossing the 0.10 threshold. That pass is not a population false-positive-rate guarantee. The same evidence thus leads to different operational conclusions when validity requirements are applied. We do not claim channel capacity, the absence of an untested channel, or learned-monitor superiority.