paper-with-me

Papers

Matrix dissimilarities based on differences in moments and sparsity

2024-06-04 · Li Tuobang

Generating a dissimilarity matrix is typically the first step in big data analysis. Although numerous methods exist, such as Euclidean distance, Minkowski distance, Manhattan distance, Bray Curtis dissimilarity, Jaccard similarity and Dice dissimilarity, it remains unclear which factors drive dissimilarity between groups. In this paper, we introduce an approach based on differences in moments and sparsity. We show that this method can delineate the key factors underlying group differences. For example, in biology, mean dissimilarity indicates differences driven by up down regulated gene expressions, standard deviation dissimilarity reflects the heterogeneity of response to treatment, and sparsity dissimilarity corresponds to differences prompted by the activation silence of genes. Through extensive reanalysis of genome, transcriptome, proteome, metabolome, immune profiling, microbiome, and social science datasets, we demonstrate insights not captured in previous studies. For instance, it shows that the sparsity dissimilarity is as effective as the mean dissimilarity in predicting the alleviation effects of a COVID 19 drug, suggesting that sparsity dissimilarity is highly meaningful.

📄 PDF Abstract BibTeX arXiv:2406.02051

Code (1)

tubanlee/MD 공식 구현

Similar Papers 제목 키워드 기반

Robust Principal Component Analysis with Non-Sparse Errors

2019-11-13

We show that when a high-dimensional data matrix is the sum of a low-rank matrix and a random error matrix with independent entries, the low-rank component can be consistently estimated by solving a convex minimization p…

Finding Exemplars from Pairwise Dissimilarities via Simultaneous Sparse Recovery

2012-12-01 · NeurIPS 2012 12 · Ehsan Elhamifar, Guillermo Sapiro, René Vidal

Given pairwise dissimilarities between data points, we consider the problem of finding a subset of data points called representatives or exemplars that can efficiently describe the data collection. We formulate the probl…

ClustGeo: an R package for hierarchical clustering with spatial constraints

2017-07-12 · Marie Chavent, Vanessa Kuentz-Simonet, Amaury Labenne, Jérôme Saracco

In this paper, we propose a Ward-like hierarchical clustering algorithm including spatial/geographical constraints. Two dissimilarity matrices $D_0$ and $D_1$ are inputted, along with a mixing parameter $\alpha \in [0,1]…

Clustering

An Automated Vehicle (AV) like Me? The Impact of Personality Similarities and Differences between Humans and AVs

2019-09-11 · Qiaoning Zhang, Connor Esterwood, X. Jessie Yang, Lionel P. Robert Jr

To better understand the impacts of similarities and dissimilarities in human and AV personalities we conducted an experimental study with 443 individuals. Generally, similarities in human and AV personalities led to a h…

Topolow: Force-Directed Euclidean Embedding of Dissimilarity Data with Robustness Against Non-Metricity and Sparsity

2025-08-03 · Omid Arhami, Pejman Rohani arxiv

The problem of embedding a set of objects into a low-dimensional Euclidean space based on a matrix of pairwise dissimilarities is fundamental in data analysis, machine learning, and statistics. However, the assumptions o…