A Genomics-Guided Multimodal Contrastive Learning Framework for Clinically Significant Prostate Cancer Risk Stratification with Missing Clinical Data.

Simple Summary: This research aims to improve how different types of medical data, including genetic information, MRI scans, and tissue images, can be combined to support accurate and reliable clinically significant prostate cancer risk stratification in patients with confirmed prostate cancer, espe...

Descripción completa

Detalles Bibliográficos
Publicado en:Cancers Vol. 18; no. 12; pp. 1952 - 1985
Autores principales: Abdullah, Shahid, Muhammad, Ather, Muhammad Ateeb, Fatima, Zulaikha, Mejorada, Carlos Guzmán Sánchez, Ruiz, Miguel Jesús Torres, Téllez, Rolando Quintero, Mata-Rivera, Miguel Félix, Zagal-Flores, Roberto
Formato: diagnostic images equations & formulas pictorial research tables/charts Journal Article
Publicado: MDPI Jun2026
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Simple Summary: This research aims to improve how different types of medical data, including genetic information, MRI scans, and tissue images, can be combined to support accurate and reliable clinically significant prostate cancer risk stratification in patients with confirmed prostate cancer, especially when some data are missing or collected from different sources. The authors develop a flexible artificial intelligence framework that uses genetic information as the main guide for learning relationships between multiple data types without requiring complete patient records. The proposed system is designed to improve low-risk versus higher-risk/csPCa discrimination, reduce critical risk-stratification errors, and remain effective in resource-limited settings where only one type of data may be available. These findings may help researchers build more robust and scalable multimodal learning systems for healthcare and other complex information domains, while improving interpretability and reliability in real-world decision-support applications. Background: Heterogeneous data integration remains a major challenge in intelligent information systems, particularly under missing-modality and cross-domain conditions. Existing multimodal fusion approaches often rely on complete datasets and weak alignment mechanisms, limiting their robustness and practical applicability. Objectives: This study aims to develop and evaluate a genomics-guided multimodal representation learning framework that enables robust heterogeneous data fusion, reliable cross-modal correspondence, and accurate prediction under incomplete-data conditions. Methods: We propose a multimodal learning architecture that models genomics as the primary biological anchor and learns conditional projections to imaging modalities, including multiparametric MRI and whole-slide histopathology (WSI). The framework formulates multimodal fusion as a genomics-guided contrastive learning problem, incorporates domain-specific optimization constraints, and learns a latent shared-state representation to support inference without requiring fully paired datasets. Evaluation was conducted using public datasets, including TCGA-PRAD and TCIA, across low-risk versus higher-risk/clinically significant prostate cancer (csPCa) discrimination, Gleason-based risk stratification, and clinically significant outcome prediction tasks under realistic multimodal and missing-modality scenarios. Results: In the adequately powered G e n o m i c s + W S I cohort (n = 486), the framework achieved an AUROC of 0.985 ± 0.005 for low-risk versus higher-risk/csPCa discrimination (p < 0.001). Exploratory analysis in a small, matched G e n o m i c s + M R I cohort (n = 28) yielded an AUROC of 0.980 ± 0.006 for the same endpoint; these findings are reported descriptively with bootstrap confidence intervals due to limited sample size. Because the negative reference group consisted of low-risk prostate cancer cases rather than cancer-free controls, results are interpreted as within-cancer risk discrimination rather than de novo cancer detection. The framework achieved weighted accuracy up to 92.1%, Cohen's κ up to 0.86, and reduced critical decision errors by 58%. Calibration remained strong (ECE 0.021–0.024), and decision-curve analysis indicated improved utility with reduced unnecessary invasive workups in retrospective modeling. Robustness analysis demonstrated AUROC degradation below 0.04 under domain shifts. Single-modality inference using genomics alone maintained AUROC > 0.90. Interpretability analysis revealed feature attributions aligned with domain-relevant genomic markers. Conclusions: The proposed framework provides a scalable and generalizable solution for heterogeneous multimodal data fusion, supporting reliable prediction, robustness to missing modalities, and applicability to complex information systems beyond the studied domain.