paper-with-me

Papers

Bias Discovery within Human Raters: A Case Study of the Jigsaw Dataset

2022-06-01 · NLPerspectives (LREC) 2022 6 · Marta Marchiori Manerba, Riccardo Guidotti, Lucia Passaro, Salvatore Ruggieri

Understanding and quantifying the bias introduced by human annotation of data is a crucial problem for trustworthy supervised learning. Recently, a perspectivist trend has emerged in the NLP community, focusing on the inadequacy of previous aggregation schemes, which suppose the existence of single ground truth. This assumption is particularly problematic for sensitive tasks involving subjective human judgments, such as toxicity detection. To address these issues, we propose a preliminary approach for bias discovery within human raters by exploring individual ratings for specific sensitive topics annotated in the texts. Our analysis’s object consists of the Jigsaw dataset, a collection of comments aiming at challenging online toxicity identification.

📄 PDF Abstract BibTeX

Code (1)

martamarchiori/bias-discovery-in-human-raters 공식 구현

Similar Papers 제목 키워드 기반

ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation

2024-05-08 · Ana Brassard, Benjamin Heinzerling, Keito Kudo, Keisuke Sakaguchi 외

Evaluating the quality of free-text explanations is a multifaceted, subjective, and labor-intensive task. Large language models (LLMs) present an appealing alternative due to their potential for consistency, scalability,…

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation

2026-04-25 · Yizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi Wang arxiv

Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks. We propose ProEval, a proactive evaluation framework t…

Gaussian ProcessesTransfer Learning

Bayesian Prediction-Powered Inference

2024-05-09 · R. Alex Hofer, Joshua Maynez, Bhuwan Dhingra, Adam Fisch 외

Prediction-powered inference (PPI) is a method that improves statistical estimates based on limited human-labeled data. Specifically, PPI methods provide tighter confidence intervals by combining small amounts of human-l…

Bayesian InferencePrediction

The Effect of Idea Elaboration on the Automatic Assessment of Idea Originality

2026-04-22 · Umberto Domanti, Moritz Mock, Sergio Agnoli, Antonella De Angeli arxiv

Automatic systems are increasingly used to assess the originality of responses in creative tasks. They offer a potential solution to key limitations of human assessment (cost, fatigue, and subjectivity), but there is pre…

Human Bias in the Face of AI: The Role of Human Judgement in AI Generated Text Evaluation

2024-09-29 · Tiffany Zhu, Iain Weissburg, Kexun Zhang, William Yang Wang

As AI advances in text generation, human trust in AI generated content remains constrained by biases that go beyond concerns of accuracy. This study explores how bias shapes the perception of AI versus human generated co…

Text Generation