Multi-way Clustering and Discordance Analysis through Deep Collective Matrix Tri-Factorization
Heterogeneous multi-typed, multimodal relational data is increasingly available in many domains and their exploratory analysis poses several challenges. We advance the state-of-the-art in neural unsupervised learning to analyze such data. We design the first neural method for collective matrix tri-factorization of arbitrary collections of matrices to perform spectral clustering of all constituent entities and learn cluster associations. Experiments on benchmark datasets demonstrate its efficacy over previous non-neural approaches. Leveraging signals from multi-way clustering and collective matrix completion we design a unique technique, called Discordance Analysis, to reveal information discrepancies across subsets of matrices in a collection with respect to two entities. We illustrate its utility in quality assessment of knowledge bases and in improving representation learning.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringMatrix CompletionRepresentation LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Discordance-Aware Multimodal Framework with Multi-Agent Clinical Reasoning
Knee osteoarthritis frequently exhibits discordance between structural damage observed in imaging and patient-reported symptoms such as pain. This mismatch complicates clinical interpretation and patient stratification a…
Multi-way Spectral Clustering of Augmented Multi-view Data through Deep Collective Matrix Tri-factorization
We present the first deep learning based architecture for collective matrix tri-factorization (DCMTF) of arbitrary collections of matrices, also known as augmented multi-view data. DCMTF can be used for multi-way spectra…
ClusteringDiscordance Minimization-based Imputation Algorithms for Missing Values in Rating Data
Ratings are frequently used to evaluate and compare subjects in various applications, from education to healthcare, because ratings provide succinct yet credible measures for comparing subjects. However, when multiple ra…
ImputationMissing ValuesLarge Language Model-Assisted Cleaning of Report-Derived Labels in a Large-Scale Chest CT Dataset
Purpose: To evaluate whether large language model (LLM)-assisted label cleaning can identify label-report discordance in CT-RATE, a large-scale public chest CT dataset. Materials and Methods: After report-level deduplica…
Auto-encoding GPS data to reveal individual and collective behaviour
We propose an innovative and generic methodology to analyse individual and collective behaviour through individual trajectory data. The work is motivated by the analysis of GPS trajectories of fishing vessels collected f…
ManagementStochastic Block Model