paper-with-me

홈 › Papers

Correcting Sociodemographic Selection Biases for Population Prediction from Social Media

2019-11-10 · Salvatore Giorgi, Veronica Lynn, Keshav Gupta, Farhan Ahmed, Sandra Matz, Lyle Ungar, H. Andrew Schwartz

Social media is increasingly used for large-scale population predictions, such as estimating community health statistics. However, social media users are not typically a representative sample of the intended population -- a "selection bias". Within the social sciences, such a bias is typically addressed with restratification techniques, where observations are reweighted according to how under- or over-sampled their socio-demographic groups are. Yet, restratifaction is rarely evaluated for improving prediction. In this two-part study, we first evaluate standard, "out-of-the-box" restratification techniques, finding they provide no improvement and often even degraded prediction accuracies across four tasks of esimating U.S. county population health statistics from Twitter. The core reasons for degraded performance seem to be tied to their reliance on either sparse or shrunken estimates of each population's socio-demographics. In the second part of our study, we develop and evaluate Robust Poststratification, which consists of three methods to address these problems: (1) estimator redistribution to account for shrinking, as well as (2) adaptive binning and (3) informed smoothing to handle sparse socio-demographic estimates. We show that each of these methods leads to significant improvement in prediction accuracies over the standard restratification approaches. Taken together, Robust Poststratification enables state-of-the-art prediction accuracies, yielding a 53.0% increase in variance explained (R^2) in the case of surveyed life satisfaction, and a 17.8% average increase across all tasks.

📄 PDF Abstract BibTeX arXiv:1911.03855

Code (1)

wwbp/robust-poststratification 공식 구현

Tasks

PredictionSelection bias

Similar Papers 제목 키워드 기반

Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs

2023-11-16 · Huaman Sun, Jiaxin Pei, MinJe Choi, David Jurgens

Human judgments are inherently subjective and are actively affected by personal traits such as gender and ethnicity. While Large Language Models (LLMs) are widely used to simulate human responses across diverse contexts,…

Identification of Bias Against People with Disabilities in Sentiment Analysis and Toxicity Detection Models

2021-11-25 · Pranav Narayanan Venkit, Shomir Wilson

Sociodemographic biases are a common problem for natural language processing, affecting the fairness and integrity of its applications. Within sentiment analysis, these biases may undermine sentiment predictions for text…

FairnessSentiment AnalysisToxic Comment Classification

Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models

2025-01-29 · YuXuan Li, Hirokazu Shirado, Sauvik Das

While advances in fairness and alignment have helped mitigate overt biases exhibited by large language models (LLMs) when explicitly prompted, we hypothesize that these models may still exhibit implicit biases when simul…

Decision MakingFairness

Sociodemographic Biases in Educational Counselling by Large Language Models

2026-04-03 · Tomasz Adamczyk, Wiktoria Mieleszczenko-Kowszewicz, Beata Bajcar, Grzegorz Chodak 외 arxiv

As Large Language Models (LLMs) are increasingly integrated into educational settings, understanding their potential biases is critical. This study examines sociodemographic biases in LLM-based educational counselling. W…

Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare

2025-02-18 · Hiba Ahsan, Arnab Sen Sharma, Silvio Amir, David Bau 외

We know from prior work that LLMs encode social biases, and that this manifests in clinical tasks. In this work we adopt tools from mechanistic interpretability to unveil sociodemographic representations and biases withi…