Placement, Not Coverage: Position-Aware CRF Emissions for Proteolytic Segmentation
Igor Shabarov ⋅ Tatiana Shashkova
Abstract
Predicting where proteases cut a precursor protein is a typed segmentation problem, and it sits under peptide therapeutics, vaccine design and proteomics. Its benchmark counts a segment as correct when both of its ends land within $\pm3$ residues of a true one. Segments are only 5 to 50 residues long, so that window rewards finding segments over placing them, and a model that places them better gains almost nothing for it. We target placement. The strongest model for this formulation decodes with a conditional random field that distinguishes a segment's start, interior and end but feeds all of those states one emission per label. We add a zero-initialized boundary head that gives those states position-specific evidence, and a lightweight adapter re-projecting the frozen protein language model embedding before the encoder. Under the $5\times4$ nested cross-validation the baseline uses, on a rebuilt UniProtKB/Swiss-Prot 2026 dataset, they add $0.054$ segment F1 over a $0.576$ base and reach $0.666$ on ESM-C 6B. What they buy is mostly placement rather than coverage: with detection held fixed, the share of boundaries landing on the annotated residue rises from $0.62$ to $0.70$ on all five outer folds. The two act on different errors, the head a precision effect and the adapter a recall effect, and the head is worth two and a half times more on the wider embedding, so capacity the base decoder cannot use is capacity the additions can. We also quantify why effects of this size are hard to measure here at all, since homology-aware folds are not exchangeable and their spread exceeds either addition alone.
Chat is not available.
Successful Page Load