paper-with-me

Papers

Filtering Context Mitigates Scarcity and Selection Bias in Political Ideology Prediction

2023-02-01 · Chen Chen, Dylan Walker, Venkatesh Saligrama

We propose a novel supervised learning approach for political ideology prediction (PIP) that is capable of predicting out-of-distribution inputs. This problem is motivated by the fact that manual data-labeling is expensive, while self-reported labels are often scarce and exhibit significant selection bias. We propose a novel statistical model that decomposes the document embeddings into a linear superposition of two vectors; a latent neutral \emph{context} vector independent of ideology, and a latent \emph{position} vector aligned with ideology. We train an end-to-end model that has intermediate contextual and positional vectors as outputs. At deployment time, our model predicts labels for input documents by exclusively leveraging the predicted positional vectors. On two benchmark datasets we show that our model is capable of outputting predictions even when trained with as little as 5\% biased data, and is significantly more accurate than the state-of-the-art. Through crowd-sourcing we validate the neutrality of contextual vectors, and show that context filtering results in ideological concentration, allowing for prediction on out-of-distribution examples.

📄 PDF Abstract BibTeX arXiv:2302.00239

Code (0)

등록된 구현이 없습니다.

Tasks

Selection bias

Similar Papers 제목 키워드 기반

Context-Aware Counterfactual Data Augmentation for Gender Bias Mitigation in Language Models

2026-02-10 · Shweta Parihar, Liu Guangliang, Natalie Parde, Lu Cheng arxiv

A challenge in mitigating social bias in fine-tuned language models (LMs) is the potential reduction in language modeling capability, which can harm downstream performance. Counterfactual data augmentation (CDA), a widel…

Data Augmentation

Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models

2025-09-20 · Md. Atabuzzaman, Ali Asgarov, Chris Thomas arxiv

Large Vision-Language Models (LVLMs) have achieved strong performance on vision-language tasks, particularly Visual Question Answering (VQA). While prior work has explored unimodal biases in VQA, the problem of selection…

Visual Question AnsweringSemantic SimilarityVisual Reasoning

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

2026-01-06 · Yusheng Dai, Zehua Chen, Yuxuan Jiang, Baolong Gao 외 arxiv

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges…

Audio Generation

Enhanced Gene Selection in Single-Cell Genomics: Pre-Filtering Synergy and Reinforced Optimization

2024-06-11 · Weiliang Zhang, Zhen Meng, Dongjie Wang, Min Wu 외

Recent advancements in single-cell genomics necessitate precision in gene panel selection to interpret complex biological data effectively. Those methods aim to streamline the analysis of scRNA-seq data by focusing on th…

Reinforcement Learning (RL)

Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes

2024-07-04 · Yusuke Hirota, Jerone T. A. Andrews, Dora Zhao, Orestis Papakyriakopoulos 외

We tackle societal bias in image-text datasets by removing spurious correlations between protected groups and image attributes. Traditional methods only target labeled attributes, ignoring biases from unlabeled ones. Usi…

Image Captioningimage-classificationImage ClassificationMulti-Label Image Classification