Query–Key Ensembles for Mathematical Reasoning Verification
Abstract
Verifying a generated mathematical solution requires a reliable correctness signal. Query-key alignment offers an internal scoring signal, but existing readouts rely on labeled selection of individual attention heads. We introduce two ensemble methods that combine information across heads without updating the model. CF-QK weights semantic verdict preferences by their agreement across token positions and responsiveness across an unlabeled batch, eliminating labeled calibration. SA-QK uses simulated annealing to select a compact ensemble when calibration labels are available. On Qwen3-8B/MATH-500, CF-QK improves balanced accuracy from 59.00\% to 74.05\% over verdict logits; SA-QK reaches 80.72\% and transfers to GSM8K and SVAMP without target labels. Experiments across model families, readout positions, and controlled solution perturbations characterize the conditions under which these gains persist. Activation patching further links selected query-key components to the model's verdict scores. Together, the results establish head aggregation as a practical approach to internal verification and identify batch composition and distribution shift as central challenges.