Ranking Is Not an Alarm: Testing Deception-Monitor Thresholds and Jacobian-Space Readouts Across Tasks, Prefixes, and Prompts
Abstract
Runtime monitors can halt or review a model response before it leads to an external action. Such a monitor needs a threshold that remains useful on a new task, during a growing response, and after a change to the surrounding prompt. Most evaluations instead report how completed responses rank. We audit the 20 released Black-to-White deception monitors whose benign thresholds can be reconstructed. None meets a fixed operating-point criterion across the target tasks: most exceed the false-alarm limit, while the three variants below it make no observed Backdoors detections and remain statistically inconclusive. This failure persists on observed rates throughout the tested false-alarm budgets and detection floors. Per-task recalibration with target labels recovers detection for the Pairs probes, tracing the main failure to threshold transfer rather than score ordering. Token-level results reveal a second problem. Mean pooling must be calibrated over complete benign prefix trajectories, whereas a final peak score already contains the largest score seen online. Because B2W does not release activations for matched prompt interventions, we examine prompt context with a separate hidden-state readout. Under the selected configuration, a matched system-persona change moves its score more than the descriptive target-class contrast; a direct residual readout behaves similarly. Deployment evaluations should follow the alarm threshold across tasks and response prefixes and should include prompt changes that leave the completion fixed.