paper-with-me

Papers

Using Simpson's Paradox to Discover Interesting Patterns in Behavioral Data

2018-05-08 · Nazanin Alipourfard, Peter G. Fennell, Kristina Lerman

We describe a data-driven discovery method that leverages Simpson's paradox to uncover interesting patterns in behavioral data. Our method systematically disaggregates data to identify subgroups within a population whose behavior deviates significantly from the rest of the population. Given an outcome of interest and a set of covariates, the method follows three steps. First, it disaggregates data into subgroups, by conditioning on a particular covariate, so as minimize the variation of the outcome within the subgroups. Next, it models the outcome as a linear function of another covariate, both in the subgroups and in the aggregate data. Finally, it compares trends to identify disaggregations that produce subgroups with different behaviors from the aggregate. We illustrate the method by applying it to three real-world behavioral datasets, including Q\&A site Stack Exchange and online learning platforms Khan Academy and Duolingo.

📄 PDF Abstract BibTeX arXiv:1805.03094

Code (1)

ninoch/Trend-Simpsons-Paradox 공식 구현

Similar Papers 제목 키워드 기반

Simpson's Paradox in Behavioral Curves: How Aggregation Distorts Parametric Models of User Dynamics

2026-05-10 · Chao Zhou arxiv

Behavioral curve modeling -- fitting parametric functions to engagement-versus-exposure data -- is standard practice in recommendation, advertising, and clinical dosing. We show that aggregation introduces a systematic d…

Classifier calibration

Resolution of Simpson's paradox via the common cause principle

2024-03-01 · A. Hovhannisyan, A. E. Allahverdyan

Simpson's paradox is an obstacle to establishing a probabilistic association between two events $a_1$ and $a_2$, given the third (lurking) random variable $B$. We focus on scenarios when the random variables $A$ (which c…

valid

Causal Collaborative Filtering

2021-02-03 · Shuyuan Xu, Yingqiang Ge, Yunqi Li, Zuohui Fu 외

Many of the traditional recommendation algorithms are designed based on the fundamental idea of mining or learning correlative patterns from data to estimate the user-item correlative preference. However, pure correlativ…

Collaborative FilteringcounterfactualRecommendation Systems

De-paradox Tree: Breaking Down Simpson's Paradox via A Kernel-Based Partition Algorithm

2026-03-02 · Xian Teng, Yu-Ru Lin arxiv

Real-world observational datasets and machine learning have revolutionized data-driven decision-making, yet many models rely on empirical associations that may be misleading due to confounding and subgroup heterogeneity.…

Causal Inference

Omitted Labels Induce Nontransitive Paradoxes in Causality

2023-11-12 · Bijan Mazaheri, Siddharth Jain, Matthew Cook, Jehoshua Bruck

We explore "omitted label contexts," in which training data is limited to a subset of the possible labels. This setting is standard among specialized human experts or specific, focused studies. By studying Simpson's para…

Causal Inference