paper-with-me

홈 › Papers

"You are an expert annotator": Automatic Best-Worst-Scaling Annotations for Emotion Intensity Modeling

2024-03-26 · Christopher Bagdon, Prathamesh Karmalker, Harsha Gurulingappa, Roman Klinger

Labeling corpora constitutes a bottleneck to create models for new tasks or domains. Large language models mitigate the issue with automatic corpus labeling methods, particularly for categorical annotations. Some NLP tasks such as emotion intensity prediction, however, require text regression, but there is no work on automating annotations for continuous label assignments. Regression is considered more challenging than classification: The fact that humans perform worse when tasked to choose values from a rating scale lead to comparative annotation methods, including best-worst scaling. This raises the question if large language model-based annotation methods show similar patterns, namely that they perform worse on rating scale annotation tasks than on comparative annotation tasks. To study this, we automate emotion intensity predictions and compare direct rating scale predictions, pairwise comparisons and best-worst scaling. We find that the latter shows the highest reliability. A transformer regressor fine-tuned on these data performs nearly on par with a model trained on the original manual annotations.

📄 PDF Abstract BibTeX arXiv:2403.17612

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Modelregression

Similar Papers 제목 키워드 기반

Best-Worst Scaling More Reliable than Rating Scales: A Case Study on Sentiment Intensity Annotation

2017-12-05 · ACL 2017 7 · Svetlana Kiritchenko, Saif M. Mohammad

Rating scales are a widely used method for data annotation; however, they present several challenges, such as difficulty in maintaining inter- and intra-annotator consistency. Best-worst scaling (BWS) is an alternative m…

Capturing Reliable Fine-Grained Sentiment Associations by Crowdsourcing and Best-Worst Scaling

2017-12-05 · Svetlana Kiritchenko, Saif M. Mohammad

Access to word-sentiment associations is useful for many applications, including sentiment analysis, stance detection, and linguistic analysis. However, manually assigning fine-grained sentiment association scores to wor…

Sentiment AnalysisStance Detection

Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?

2026-05-12 · Yichen Feng, Yuetai Li, Chunjiang Liu, Yuanyuan Chen 외 arxiv

Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an explicit aesthetic judgment. Most existing s…

At the Lower End of Language---Exploring the Vulgar and Obscene Side of German

2019-08-01 · WS 2019 8 · Elisabeth Eder, Ulrike Krieg-Holz, Udo Hahn

In this paper, we describe a workflow for the data-driven acquisition and semantic scaling of a lexicon that covers lexical items from the lower end of the German language register{---}terms typically considered as rough…

Tell Me What You Read: Automatic Expertise-Based Annotator Assignment for Text Annotation in Expert Domains

2021-09-01 · RANLP 2021 9 · Hiyori Yoshikawa, Tomoya Iwakura, Kimi Kaneko, Hiroaki Yoshida 외

This paper investigates the effectiveness of automatic annotator assignment for text annotation in expert domains. In the task of creating high-quality annotated corpora, expert domains often cover multiple sub-domains (…

text annotation