paper-with-me

홈 › Papers

A Framework to Adjust Dependency Measure Estimates for Chance

2015-10-27 · Simone Romano, Nguyen Xuan Vinh, James Bailey, Karin Verspoor

Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the Maximal Information Coefficient (MIC); and categorical variables that are dependent to the target class are selected using Gini gain in random forests. Nonetheless, because dependency measures are estimated on finite samples, the interpretability of their quantification and the accuracy when ranking dependencies become challenging. Dependency estimates are not equal to 0 when variables are independent, cannot be compared if computed on different sample size, and they are inflated by chance on variables with more categories. In this paper, we propose a framework to adjust dependency measure estimates on finite samples. Our adjustments, which are simple and applicable to any dependency measure, are helpful in improving interpretability when quantifying dependency and in improving accuracy on the task of ranking dependencies. In particular, we demonstrate that our approach enhances the interpretability of MIC when used as a proxy for the amount of noise between variables, and to gain accuracy when ranking variables during the splitting procedure in random forests.

📄 PDF Abstract BibTeX arXiv:1510.07786

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

Establishing Annotation Quality in Multi-label Annotations

2022-10-01 · COLING 2022 10 · Marian Marchal, Merel Scholman, Frances Yung, Vera Demberg

In many linguistic fields requiring annotated data, multiple interpretations of a single item are possible. Multi-label annotations more accurately reflect this possibility. However, allowing for multi-label annotations …

Adjusting for Chance Clustering Comparison Measures

2015-12-03 · Simone Romano, Nguyen Xuan Vinh, James Bailey, Karin Verspoor

Adjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on…

Clustering

Two-Stage TMLE to Reduce Bias and Improve Efficiency in Cluster Randomized Trials

2021-06-29 · Laura B. Balzer, Mark van der Laan, James Ayieko, Moses Kamya 외

Cluster randomized trials (CRTs) randomly assign an intervention to groups of individuals (e.g., clinics or communities) and measure outcomes on individuals in those groups. While offering many advantages, this experimen…

Experimental Design

Selection Bias Correction and Effect Size Estimation under Dependence

2014-05-16 · Kean Ming Tan, Noah Simon, Daniela Witten

We consider large-scale studies in which it is of interest to test a very large number of hypotheses, and then to estimate the effect sizes corresponding to the rejected hypotheses. For instance, this setting arises in t…

Selection bias

Multivariate Dependency Measure based on Copula and Gaussian Kernel

2017-08-24 · Angshuman Roy, Alok Goswami, C. A. Murthy

We propose a new multivariate dependency measure. It is obtained by considering a Gaussian kernel based distance between the copula transform of the given d-dimensional distribution and the uniform copula and then approp…