The identification of prognostic biomarkers from gene expression data represents a key challenge in modern biomedical research, specially in the context of survival analysis. A crucial step is the quantification of distributional differences in gene expression between clinically distinct patient groups, such as high-and low-risk cohorts. In this work, we employ the Jensen–Shannon Divergence (JSD) as a measure of dissimilarity between empirical gene expression distributions. Unlike classical measures, the JSD is bounded, robust to sparse data, and directly interpretable, making it especially suitable for high-dimensional omic datasets. To ensure robustness, we complement JSD estimation with a procedure that allows the assessment of statistical variability and the derivation of empirical p -values. Genes exhibiting high JSD scores and significantly low p-values are considered potential prognostic biomarkers, as their expression patterns display strong divergence between patient populations. This framework provides a statistically rigorous and interpretable approach to biomarker discovery, integrating information-theoretic distance measures with permutation testing approach in survival-oriented genomic analysis.
A Gene-Wise Jensen-Shannon Divergence Framework for Prognostic Biomarker Discovery
Di Crescenzo, Antonio;Iuliano, Antonella
2026
Abstract
The identification of prognostic biomarkers from gene expression data represents a key challenge in modern biomedical research, specially in the context of survival analysis. A crucial step is the quantification of distributional differences in gene expression between clinically distinct patient groups, such as high-and low-risk cohorts. In this work, we employ the Jensen–Shannon Divergence (JSD) as a measure of dissimilarity between empirical gene expression distributions. Unlike classical measures, the JSD is bounded, robust to sparse data, and directly interpretable, making it especially suitable for high-dimensional omic datasets. To ensure robustness, we complement JSD estimation with a procedure that allows the assessment of statistical variability and the derivation of empirical p -values. Genes exhibiting high JSD scores and significantly low p-values are considered potential prognostic biomarkers, as their expression patterns display strong divergence between patient populations. This framework provides a statistically rigorous and interpretable approach to biomarker discovery, integrating information-theoretic distance measures with permutation testing approach in survival-oriented genomic analysis.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


