Characterization and reduction of variability in selection based on effect-size using association measures in cohort study of heterogeneous diseases
Cohort studies employ pairwise measures of association to quantify dependencies among conditions and exposures. To reliably use these measures to draw conclusions about the underlying association strengths requires that the measures be robust and unbiased. These considerations assume greater significance when applied to disease networks, where associations among heterogeneous pairs of diseases are ranked. Using disease diagnoses data from a large cohort of 5.5 million individuals, we develop a comprehensive methodology to characterize the bias of standard association measures like relative risk and $\phi$ correlation. To overcome these biases, we devise a novel measure based on a stochastic model for disease development. The new measure is demonstrated to have the least overall bias and hence would be most suitable for application to heterogeneous disease cohorts.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Impact of Coreset Selection on Spurious Correlations and Group Robustness
Coreset selection methods have shown promise in reducing the training data size while maintaining model performance for data-efficient machine learning. However, as many datasets suffer from biases that cause models to l…
Variance-Reduced Heterogeneous Federated Learning via Stratified Client Selection
Client selection strategies are widely adopted to handle the communication-efficient problem in recent studies of Federated Learning (FL). However, due to the large variance of the selected subset's update, prior selecti…
DiversityFederated LearningUnsupervised Feature Selection Based on the Morisita Estimator of Intrinsic Dimension
This paper deals with a new filter algorithm for selecting the smallest subset of features carrying all the information content of a data set (i.e. for removing redundant features). It is an advanced version of the fract…
Dimensionality Reductionfeature selectionUnsupervised Bump Hunting Using Principal Components
Principal Components Analysis is a widely used technique for dimension reduction and characterization of variability in multivariate populations. Our interest lies in studying when and why the rotation to principal compo…
Dimensionality ReductionAre we describing the same sound? An analysis of word embedding spaces of expressive piano performance
Semantic embeddings play a crucial role in natural language-based information retrieval. Embedding models represent words and contexts as vectors whose spatial configuration is derived from the distribution of words in l…
Information RetrievalRetrievalSemantic SimilaritySemantic Textual Similarity