paper-with-me

홈 › Papers

Comparative Evaluation of Label-Agnostic Selection Bias in Multilingual Hate Speech Datasets

2020-11-01 · EMNLP 2020 11 · Nedjma Ousidhoum, Yangqiu Song, Dit-yan Yeung

Work on bias in hate speech typically aims to improve classification performance while relatively overlooking the quality of the data. We examine selection bias in hate speech in a language and label independent fashion. We first use topic models to discover latent semantics in eleven hate speech corpora, then, we present two bias evaluation metrics based on the semantic similarity between topics and search words frequently used to build corpora. We discuss the possibility of revising the data collection process by comparing datasets and analyzing contrastive case studies.

📄 PDF Abstract BibTeX

Code (1)

HKUST-KnowComp/HS_Bias_Eval

Tasks

Hate Speech DetectionSelection biasSemantic SimilaritySemantic Textual SimilarityTopic Models

Similar Papers 제목 키워드 기반

Debiased Graph Neural Networks with Agnostic Label Selection Bias

2022-01-19 · Shaohua Fan, Xiao Wang, Chuan Shi, Kun Kuang 외

Most existing Graph Neural Networks (GNNs) are proposed without considering the selection bias in data, i.e., the inconsistent distribution between the training set with test set. In reality, the test data is not even av…

parameter estimationSelection bias

Navigating Towards Fairness with Data Selection

2024-12-15 · Yixuan Zhang, Zhidong Li, Yang Wang, Fang Chen 외

Machine learning algorithms often struggle to eliminate inherent data biases, particularly those arising from unreliable labels, which poses a significant challenge in ensuring fairness. Existing fairness techniques that…

FairnessHoldout Set

Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs

2026-05-28 · Mahjabin Nahar, Nafis Irtiza Tripto, Aiping Xiong, Ting-Hao 'Kenneth' Huang 외 arxiv

As AI-generated and AI-assisted content floods online spaces, source labels attached to such content can distort human reasoning judgments, with downstream consequences for moderation, evaluation, and decision-making. Wh…

Logical Fallacies

Knowledge-Guided Biomarker Identification for Label-Free Single-Cell RNA-Seq Data: A Reinforcement Learning Perspective

2025-01-02 · Meng Xiao, Weiliang Zhang, Xiaohan Huang, HengShu Zhu 외

Gene panel selection aims to identify the most informative genomic biomarkers in label-free genomic datasets. Traditional approaches, which rely on domain expertise, embedded machine learning models, or heuristic-based i…

No evaluation without fair representation : Impact of label and selection bias on the evaluation, performance and mitigation of classification models

2026-03-10 · Magali Legast, Toon Calders, François Fouss arxiv

Bias can be introduced in diverse ways in machine learning datasets, for example via selection or label bias. Although these bias types in themselves have an influence on important aspects of fair machine learning, their…