paper-with-me

Papers

Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

2023-11-01 · Senjuti Dutta, Sid Mittal, Sherol Chen, Deepak Ramachandran, Ravi Rajakumar, Ian Kivlichan, Sunny Mak, Alena Butryna, Praveen Paritosh

The prevalence and impact of toxic discussions online have made content moderation crucial.Automated systems can play a vital role in identifying toxicity, and reducing the reliance on human moderation.Nevertheless, identifying toxic comments for diverse communities continues to present challenges that are addressed in this paper.The two-part goal of this study is to(1)identify intuitive variances from annotator disagreement using quantitative analysis and (2)model the subjectivity of these viewpoints.To achieve our goal, we published a new dataset\footnote{\url{https://github.com/XXX}} with expert annotators' annotations and used two other public datasets to identify the subjectivity of toxicity.Then leveraging the Large Language Model(LLM),we evaluate the model's ability to mimic diverse viewpoints on toxicity by varying size of the training data and utilizing same set of annotators as the test set used during model training and a separate set of annotators as the test set.We conclude that subjectivity is evident across all annotator groups, demonstrating the shortcomings of majority-rule voting. Moving forward, subjective annotations should serve as ground truth labels for training models for domains like toxicity in diverse communities.

📄 PDF Abstract BibTeX arXiv:2311.00203

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Chord Label Personalization through Deep Learning of Integrated Harmonic Interval-based Representations

2017-06-29 · H. V. Koops, W. B. de Haas, J. Bransen, A. Volk

The increasing accuracy of automatic chord estimation systems, the availability of vast amounts of heterogeneous reference annotations, and insights from annotator subjectivity research make chord label personalization i…

You Are What You Annotate: Towards Better Models through Annotator Representations

2023-05-24 · Naihao Deng, Xinliang Frederick Zhang, Siyang Liu, Winston Wu 외

Annotator disagreement is ubiquitous in natural language processing (NLP) tasks. There are multiple reasons for such disagreements, including the subjectivity of the task, difficult cases, unclear guidelines, and so on. …

Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale

2025-12-10 · Karl Gustav Gailit, Kadri Muischnek, Kairit Sirts arxiv

This article presents the creation of an Estonian-language dataset for document-level subjectivity, analyzes the resulting annotations, and reports an initial experiment of automatic subjectivity analysis using a large l…

Subjectivity Analysis

Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments

2025-09-08 · Amir Homayounirad, Enrico Liscio, Tong Wang, Catholijn M. Jonker 외 arxiv

Aggregating multiple annotations into a single ground truth label may hide valuable insights into annotator disagreement, particularly in tasks where subjectivity plays a crucial role. In this work, we explore methods fo…

Value prediction

SafeWebUH at SemEval-2023 Task 11: Learning Annotator Disagreement in Derogatory Text: Comparison of Direct Training vs Aggregation

2023-05-01 · Sadat Shahriar, Thamar Solorio

Subjectivity and difference of opinion are key social phenomena, and it is crucial to take these into account in the annotation and detection process of derogatory textual content. In this paper, we use four datasets pro…