paper-with-me

Papers

Corpus Considerations for Annotator Modeling and Scaling

2024-04-02 · Olufunke O. Sarumi, Béla Neuendorf, Joan Plepi, Lucie Flek, Jörg Schlötterer, Charles Welch

Recent trends in natural language processing research and annotation tasks affirm a paradigm shift from the traditional reliance on a single ground truth to a focus on individual perspectives, particularly in subjective tasks. In scenarios where annotation tasks are meant to encompass diversity, models that solely rely on the majority class labels may inadvertently disregard valuable minority perspectives. This oversight could result in the omission of crucial information and, in a broader context, risk disrupting the balance within larger ecosystems. As the landscape of annotator modeling unfolds with diverse representation techniques, it becomes imperative to investigate their effectiveness with the fine-grained features of the datasets in view. This study systematically explores various annotator modeling techniques and compares their performance across seven corpora. From our findings, we show that the commonly used user token model consistently outperforms more complex models. We introduce a composite embedding approach and show distinct differences in which model performs best as a function of the agreement with a given dataset. Our findings shed light on the relationship between corpus statistics and annotator modeling performance, which informs future work on corpus construction and perspectivist NLP.

📄 PDF Abstract BibTeX arXiv:2404.02340

Code (1)

caisa-lab/naacl2024-considerations-annotator-modeling 공식 구현 pytorch

Tasks

Diversity

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus

2026-03-17 · Martina Simonotti, Ludovica Pannitto, Eleonora Zucchini, Silvia Ballarè 외 arxiv

This paper analyses the implementation of Automatic Speech Recognition (ASR) into the transcription workflow of the KIParla corpus, a resource of spoken Italian. Through a two-phase experiment, 11 expert and novice trans…

Speech Recognition

White Paper: Challenges and Considerations for the Creation of a Large Labelled Repository of Online Videos with Questionable Content

2021-01-25 · Thamar Solorio, Mahsa Shafaei, Christos Smailis, Mona Diab 외

This white paper presents a summary of the discussions regarding critical considerations to develop an extensive repository of online videos annotated with labels indicating questionable content. The main discussion poin…

Quality Focused Approach to a Learner Corpus Development

2020-05-01 · LREC 2020 5 · Roberts Dar{\c{g}}is, Ilze Auzi{\c{n}}a, Krist{\=\i}ne Lev{\=a}ne-Petrova, Inga Kaija

The paper presents quality focused approach to a learner corpus development. The methodology was developed with multiple design considerations put in place to make the annotation process easier and at the same time reduc…

Morphological Analysis

The Measuring Hate Speech Corpus: Leveraging Rasch Measurement Theory for Data Perspectivism

2022-06-01 · NLPerspectives (LREC) 2022 6 · Pratik Sachdeva, Renata Barreto, Geoff Bacon, Alexander Sahn 외

We introduce the Measuring Hate Speech corpus, a dataset created to measure hate speech while adjusting for annotators’ perspectives. It consists of 50,070 social media comments spanning YouTube, Reddit, and Twitter, lab…

Experimental Design

Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree

2024-10-16 · Harbani Jaggi, Kashyap Murali, Eve Fleisig, Erdem Biyik

When annotators disagree, predicting the labels given by individual annotators can capture nuances overlooked by traditional label aggregation. We introduce three approaches to predicting individual annotator ratings on …

Collaborative FilteringIn-Context LearningPredictionSurvey