Why Cancer cfDNA Models Fail on Chronic Disease: A Geometric Information-Theoretic Bound and Its Architectural Implications
Fereshteh Abedini ⋅ Soheil Damangir
Abstract
Cell-free DNA (cfDNA) liquid biopsy has been clinically routine for cancer detection and monitoring for years but has yet to produce applications in chronic diseases of parenchymal organs such as fibrotic liver, kidney, and lung disease, despite diseased organs shedding orders of magnitude more cfDNA into the bloodstream than tumors do. This gap is structural: methods developed for cancer have been adapted for chronic disease without a theoretical framework establishing when this should work or where it should break down. We formalize cfDNA classification as recovering a disease-induced shift on a low-dimensional submanifold of the simplex over molecular configurations, observed through sparse genomic windows. The formalization reveals that cancer and chronic disease occupy structurally distinct regimes parameterized by one dimensionless quantity: the inverse participation ratio $c(v)$ of the disease perturbation across windows. Cancer is small-$c(v)$: signal on a few genomic hotspots. Chronic disease is large-$c(v)$: signal spread across many co-regulated regulatory programs. We prove an information-theoretic lower bound: any classifier observing $k$ windows recovers at most a $\sqrt{k/c(v)}$ fraction of the available information in the local-asymptotic-normality regime, regardless of per-cfDNA model complexity. The bound applies architecture-agnostically to classifiers that focus deep computation on a small number of pre-selected genomic regions, i.e. the architectural shape that has succeeded in cancer detection. Such classifiers degrade sharply rather than gradually as $c(v)$ grows, with the transition provable from the geometry of the problem. Chronic-disease classification requires an architecture matched to its regime: compute allocated across a very large set of windows rather than into deeper per-cfDNA models, with $k \gtrsim c(v)$. We propose a concrete architecture and verify on synthetic experiments that the regime transition predicted by the theorem is observed, and the architecture outperforms cancer-style baselines on chronic-like signal. We then validate on real patient data, where the architecture exceeds both cancer-tuned methods and standard-of-care clinical tests.
Chat is not available.
Successful Page Load