From Sparse Readout to Mechanism: An Evidence Ladder for Affective Appraisal
Abstract
Appraisal theories treat threat as a precursor of caution and blame as a precursor of accountability, so a sparse coordinate that distinguishes threatening prompts is a natural candidate mechanism for cautionary text. Interpretability can support that kind of discovery only if the candidate is tested as a mechanism rather than accepted as a readout: a coordinate that decodes an appraisal can still fail to occupy the role the theory assigns it. We ask whether sparse autoencoder (SAE) features that distinguish threat and blame in Llama-3.2-1B-Instruct also help produce cautionary and accountability-oriented responses. Features selected on 12 training pairs are tested on eight held-out variants from two unseen template families; response-side ranks, decoder-vector interventions, 864 activation patches, and a keyword behavioral test then check whether those features occupy the proposed roles. One threat feature and two blame features keep the expected direction on all eight held-out variants. Different features rank first for caution and accountability in generated text, prompt interventions leave the response features inactive at the first generated token, and the keyword changes are no larger than those from an unrelated control. The supported finding is a held-out readout, not a causal bridge from appraisal to response. We organize the tests as an evidence ladder with four rungs: held-out readout, role alignment, causal propagation, and behavioral specificity.