Cell-Type-Specific Expression Improves Survival Prediction Despite Being Hard to Predict
Gadi Naveh ⋅ Amit Miller ⋅ Gal Keinan
Abstract
Deconvolution of bulk RNA into single-cell-resolution features can boost performance in a downstream bulk-input task. Most existing methods focus on predicting the \textit{proportions} $\pi_k$ of the $K$ constituent cell types (CTs), while only a minority also predict \textit{cell-type specific expression} (CTSE) $\mu_{kg}$ in gene $g$ and CT $k$. While many methods perform well at $\pi$ prediction in typical settings, recent results show that the mean performance of $\mu$ prediction is generally much poorer. We show that this is due to information-theoretic bounds that bind certain predictor classes. Despite typically lower estimation accuracy, we provide evidence that CTSE features may have higher utility for downstream tasks, demonstrating this in overall survival (OS) prediction in non-small cell lung cancer (NSCLC). We adopt \textit{targeted deconvolution}, selecting 106 biologically motivated (CT, gene) pairs that avoid pure markers (redundant with $\pi_k$) and unrecoverable low-signal targets; variance share $\phi_{kg}$ quantifies these heuristics. We show that various deconvolved single features $\mu_{kg}$ predict OS significantly better than their bulk counterparts $x_g$ on the same gene, and that top single $\mu$ features outperform top $\pi$ features. Thus, estimation difficulty and downstream utility are largely decoupled: targets hard to estimate from bulk can still carry strong prognostic signal.
Chat is not available.
Successful Page Load