The Sensitivity of Annotator Bias to Task Definitions in Argument Mining
NLP models are dependent on the data they are trained on, including how this data is annotated. NLP research increasingly examines the social biases of models, but often in the light of their training data and specific social biases that can be identified in the text itself. In this paper, we present an annotation experiment that is the first to examine the extent to which social bias is sensitive to how data is annotated. We do so by collecting annotations of arguments in the same documents following four different guidelines and from four different demographic annotator backgrounds. We show that annotations exhibit widely different levels of group disparity depending on which guidelines annotators follow. The differences are not explained by task complexity, but rather by characteristics of these demographic groups, as previously identified by sociological studies. We release a dataset that is small in the number of instances but large in the number of annotations with demographic information, and our results encourage an increased awareness of annotator bias.
Code (1)
Tasks
Argument MiningSensitivitySimilar Papers 제목 키워드 기반
The Sensitivity of Annotator Bias to Task Definitions
NLP models are biased by the data they are trained on, including how it is annotated, but NLP research increasingly examines the em social biases of models, often in the light of their training data. This paper is first …
SensitivityA Resource for Enthymeme Detection in Controversial Political Discourse
Enthymemes, arguments with unstated premises or conclusions, are pervasive in persuasive discourse, yet their annotation remains notoriously subjective. We present a resource of 1,482 tweets from politically controversia…
Reducing annotator bias by belief elicitation
Crowdsourced annotations of data play a substantial role in the development of Artificial Intelligence (AI). It is broadly recognised that annotations of text data can contain annotator bias, where systematic disagreemen…
Are Large Language Models Reliable Argument Quality Annotators?
Evaluating the quality of arguments is a crucial aspect of any system leveraging argument mining. However, it is a challenge to obtain reliable and consistent annotations regarding argument quality, as this usually requi…
Argument MiningDefinitional Sensitivity in Media Bias Detection: A Multi-Definition Dataset and Benchmark
Media bias detection relies on definitions and examples that specify what counts as bias, yet these specifications often vary across datasets or remain implicit, even when given the same name. Such variation makes it unc…
Bias Detection