Benchmark Recovery Does Not Certify a Frozen Tool-Trigger Contract After Quantization
Abstract
Benchmark recovery is often treated as a deployment certificate for low-bit LLMs, but thresholded controller stacks ship a different object: a frozen score-to-action contract. This paper asks a narrower acceptance question: after quantization, does that interface survive independent threshold retuning against frozen FP16 labels, the strongest cheap repair an operator would plausibly try first? We study one fixed four-action controller, one frozen train/cal/report protocol, and four instruction-tuned checkpoints: Llama-3.1-8B/70B-Instruct and Qwen-2.5-7B/72B-Instruct. In the only score-carrying regime, W4A16, GPTQ + -search stays near FP16 on MMLU yet still leaves 11.9/10.9 percentage-point false tool-trigger rates and 16.1/15.4 aggregate FCR on the two large models. Because BFCL directly labels the tool frontier, that failure already establishes the paper's main acceptance claim even if FAQ and the preservation-only retrieve/stop frontiers are ignored. Aggregate FCR is reported only as interface-wide preservation context because retrieve/stop remain preservation-only. FAQ is included only as a bounded reducibility probe under the same frozen audit target, not as a generic new quantizer or the central novelty claim. FAQ lowers tool FTR and aggregate FCR without obvious MMLU or throughput collapse, but that is secondary non-collapse evidence. The contribution is therefore a single-controller deployment-acceptance counterexample under a frozen permission set: on the retained large-model W4A16 rows, benchmark-close MMLU recovery does not certify preserved tool-trigger behavior. The accompanying contribution is the freeze-and- retune audit protocol that makes that failure measurable.