When an Evaluation Should Refuse to Answer: A Fail-Closed Weight-Perturbation Pilot in Hebrew
Abstract
An evaluation can be implemented correctly and still lack the evidence needed for its headline claim. We study an evaluation of language model capabilities that must stop unless its checks for clean competence, scoring, perturbation fidelity, and dose coverage all pass. Our Modern Hebrew development pilot was precommitted and hash-bound before execution. It compared a 1.7B base checkpoint adapted to Hebrew with its 1.7B Qwen initializer, using 108 items reviewed by the author and a planned matrix of 110 Gaussian and quantization conditions. Execution ended after the exact 26-condition calibration prefix. Across both models, 108 primary and 108 linked secondary items under three prompt frames yielded 5,184 candidate realizations; every one passed strict prefix and token boundary checks. After every executed condition, the guarded model state was restored bit-for-bit, and the largest error between intended and realized dose was 0.086%. Even so, clean headroom failed in four of six model by capability cells. The capability blind PAVA retention curve ranged from 0.961 to 0.382, leaving inverse targets 0.29 and 0.15 uncovered and only two qualifying interior levels. No final grid was frozen, so no inferential or quantization condition ran. We report no capability ordering. The stop prevented 84 of 110 planned conditions (76.4%) from running after their prerequisites failed. In separately frozen analyses, 16 of 20 new calibration directions also triggered STOP. A known-truth study under a frozen scenario suite then measured the trade-off between answer coverage and unsupported risk. The case shows why reproducible execution is not enough: the evidence must also identify the claim.