Do Knowledge and Reasoning-Oriented Benchmarks Induce Stable Representation Signatures?
Parsa Rahimi ⋅ Azim Dehghani Amirabad
Abstract
Passive internal signatures could support monitoring, adaptive computation, and mechanistic analysis, but benchmark prompts can be separable without a shared computational distinction. We test whether a prespecified knowledge--reasoning label defines a reusable internal signal in five current 7B--31B models. The task is the inferential unit: probes hold out complete tasks, task labels are permuted, and the best layer--stream site is reselected for each permutation. Raw probes recover an eight-versus-eight-subject MMLU partition in all five models (macro-F1 $0.78$--$0.81$). With identical folds, dimension, classifier, regularization, and permutation null for activation and prompt features, the activation advantage is only $0.017$--$0.070$ macro-F1 and is not significant ($p=0.177$--$0.440$). At the prespecified $0.01$ level, lexical adjustment removes significance while task identity remains decodable; sites vary across panels and selectors; and two released base-training trajectories have no common direction. These pooled layerwise measures establish task-label recovery, not a transferable knowledge--reasoning signature, and do not exclude moderate effects. The resulting task-level, prompt-matched protocol provides an evaluation sequence for future internal monitors and intervention hypotheses.
Chat is not available.
Successful Page Load