paper-with-me

홈 › Papers

Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset

2024-03-28 · Janis Goldzycher, Paul Röttger, Gerold Schneider

Hate speech detection models are only as good as the data they are trained on. Datasets sourced from social media suffer from systematic gaps and biases, leading to unreliable models with simplistic decision boundaries. Adversarial datasets, collected by exploiting model weaknesses, promise to fix this problem. However, adversarial data collection can be slow and costly, and individual annotators have limited creativity. In this paper, we introduce GAHD, a new German Adversarial Hate speech Dataset comprising ca.\ 11k examples. During data collection, we explore new strategies for supporting annotators, to create more diverse adversarial examples more efficiently and provide a manual analysis of annotator disagreements for each strategy. Our experiments show that the resulting dataset is challenging even for state-of-the-art hate speech detection models, and that training on GAHD clearly improves model robustness. Further, we find that mixing multiple support strategies is most advantageous. We make GAHD publicly available at https://github.com/jagol/gahd.

📄 PDF Abstract BibTeX arXiv:2403.19559

Code (1)

jagol/gahd 공식 구현

Tasks

Hate Speech Detection

Similar Papers 제목 키워드 기반

Collecting high-quality adversarial data for machine reading comprehension tasks with humans and models in the loop

2022-06-28 · NAACL (DADC) 2022 7 · Damian Y. Romero Diaz, Magdalena Anioł, John Culnan

We present our experience as annotators in the creation of high-quality, adversarial machine-reading-comprehension data for extractive QA for Task 1 of the First Workshop on Dynamic Adversarial Data Collection (DADC). DA…

Machine Reading ComprehensionReading Comprehension

Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants

2021-12-16 · NAACL 2022 7 · Max Bartolo, Tristan Thrush, Sebastian Riedel, Pontus Stenetorp 외

In Dynamic Adversarial Data Collection (DADC), human annotators are tasked with finding examples that models struggle to predict correctly. Models trained on DADC-collected training data have been shown to be more robust…

Extractive Question-AnsweringQuestion Answering

Crowd-sourcing annotation of complex NLU tasks: A case study of argumentative content annotation

2019-11-01 · WS 2019 11 · Tamar Lavee, Lili Kotlerman, Matan Orbach, Yonatan Bilu 외

Recent advancements in machine reading and listening comprehension involve the annotation of long texts. Such tasks are typically time consuming, making crowd-annotations an attractive solution, yet their complexity ofte…

Natural Language UnderstandingReading ComprehensionSentence

Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources

2022-01-25 · Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman, Kimbo Chen 외

In recent years, large-scale data collection efforts have prioritized the amount of data collected in order to improve the modeling capabilities of large language models. This prioritization, however, has resulted in con…

The Teacher-Student Chatroom Corpus

2020-11-13 · NLP4CALL (COLING) 2020 11 · Andrew Caines, Helen Yannakoudakis, Helena Edmondson, Helen Allen 외

The Teacher-Student Chatroom Corpus (TSCC) is a collection of written conversations captured during one-to-one lessons between teachers and learners of English. The lessons took place in an online chatroom and therefore …

Descriptive