Steered affect moves a partner agent's activations, and the logistic probe that reads it returns a tie-break, not an estimate
Abstract
When one language-model agent is steered into an emotion and talks to another, a linear probe’s projection of the second agent’s reply activations moves with the dose, for five of six emotions on Qwen3.6-27B and on Llama3-8B-Instruct-abliterated. We ask what carries it, and find that the standard linear tools mislead. The logistic-regression direction that probe studies use to read such states reproduces across disjoint halves of the same pool at only 0.57 (27B) and 0.29 (8B), while difference-of-means reaches 0.97 and 0.94: with n < d every logistic fit is a perfect separator, so its direction is the penalty’s tie-break. With the reproducible direction, three pre-registered ablations test the linear account. Removing the emotion direction at every readable layer leaves a median 91% of the response; a rank-5 affect subspace fails its own instrument check; and a rank-5 subspace fitted by the same estimator on permuted labels blocks three emotions at BHFDR q < 0.05 where the affect subspace blocks two (at about half its size on those two, and more than it on afraid), and blocks angry in five and sad in four of five label permutations, while random frames of the same or half its footprint are inert. The control that bites is the one built from the pool’s activations. Across ten re-splits of the pool, only sad’s block against the permuted control survives every draw (median 35%), so the affect-specific claim rests on sad. The probe does not separate the receiver’s state from its model of the sender’s.