An interpretable multiple kernel learning approach for the discovery of integrative cancer subtypes
Due to the complexity of cancer, clustering algorithms have been used to disentangle the observed heterogeneity and identify cancer subtypes that can be treated specifically. While kernel based clustering approaches allow the use of more than one input matrix, which is an important factor when considering a multidimensional disease like cancer, the clustering results remain hard to evaluate and, in many cases, it is unclear which piece of information had which impact on the final result. In this paper, we propose an extension of multiple kernel learning clustering that enables the characterization of each identified patient cluster based on the features that had the highest impact on the result. To this end, we combine feature clustering with multiple kernel dimensionality reduction and introduce FIPPA, a score which measures the feature cluster impact on a patient cluster. Results: We applied the approach to different cancer types described by four different data types with the aim of identifying integrative patient subtypes and understanding which features were most important for their identification. Our results show that our method does not only have state-of-the-art performance according to standard measures (e.g., survival analysis), but, based on the high impact features, it also produces meaningful explanations for the molecular bases of the subtypes. This could provide an important step in the validation of potential cancer subtypes and enable the formulation of new hypotheses concerning individual patient groups. Similar analysis are possible for other disease phenotypes.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringDimensionality ReductionDiscovery Of Integrative Cancer SubtypesSurvival AnalysisSimilar Papers 제목 키워드 기반
Multiple kernel learning for integrative consensus clustering of 'omic datasets
Diverse applications - particularly in tumour subtyping - have demonstrated the importance of integrative clustering techniques for combining information from multiple data sources. Cluster-Of-Clusters Analysis (COCA) is…
ClusteringIGCN: Integrative Graph Convolution Networks for patient level insights and biomarker discovery in multi-omics integration
Developing computational tools for integrative analysis across multiple types of omics data has been of immense importance in cancer molecular biology and precision medicine research. While recent advancements have yield…
Node ClassificationMulti-Kernel LS-SVM Based Bio-Clinical Data Integration: Applications to Ovarian Cancer
The medical research facilitates to acquire a diverse type of data from the same individual for particular cancer. Recent studies show that utilizing such diverse data results in more accurate predictions. The major chal…
Data IntegrationTowards multiple kernel principal component analysis for integrative analysis of tumor samples
Personalized treatment of patients based on tissue-specific cancer subtypes has strongly increased the efficacy of the chosen therapies. Even though the amount of data measured for cancer patients has increased over the …
ClusteringData IntegrationPan-Cancer Integrative Histology-Genomic Analysis via Interpretable Multimodal Deep Learning
The rapidly emerging field of deep learning-based computational pathology has demonstrated promise in developing objective prognostic models from histology whole slide images. However, most prognostic models are either b…
Deep LearningMultimodal Deep LearningPrognosiswhole slide images