The Confessions of a Protein Language Model: Steering, Activations, and What Predicted Confidence Cannot Certify
Panini Shah ⋅ Raj Panchali ⋅ Pearl Mody ⋅ Mahek Sanghvi ⋅ MEERA NARVEKAR ⋅ Ruhina Karani
Abstract
Protein generate-and-screen pipelines use interventions and readouts to select candidates, but rarely test whether these track their intended targets. We audit published repetition steering, ESMFold confidence, and an activation readout. Steering preserves confidence at its operating point, but a fitted cliff follows at $\alpha_{50}=1.240\times$ (sweep's own control: $64.0\%$). Direction matters at moderate matched-norm push, but not at saturation. Unsteered generator activations predict the confidence gate (ROC-AUC $0.786\pm0.004$), supporting cost-aware, model-specific triage. We also replicate established length sensitivity, through which shortening can manufacture false nulls. Yet on composition- and length-matched design--scramble pairs, no balanced-accuracy interval excludes chance for experimental stability. Predicting the gate does not establish foldability. We should test what a readout measures before using it to select. Code and data: https://github.com/rpextra2026-afk/plm-steering-audit
Chat is not available.
Successful Page Load