paper-with-me

홈 › Papers

A Survey on Cost Types, Interaction Schemes, and Annotator Performance Models in Selection Algorithms for Active Learning in Classification

2021-09-23 · Marek Herde, Denis Huseljic, Bernhard Sick, Adrian Calma

Pool-based active learning (AL) aims to optimize the annotation process (i.e., labeling) as the acquisition of annotations is often time-consuming and therefore expensive. For this purpose, an AL strategy queries annotations intelligently from annotators to train a high-performance classification model at a low annotation cost. Traditional AL strategies operate in an idealized framework. They assume a single, omniscient annotator who never gets tired and charges uniformly regardless of query difficulty. However, in real-world applications, we often face human annotators, e.g., crowd or in-house workers, who make annotation mistakes and can be reluctant to respond if tired or faced with complex queries. Recently, a wide range of novel AL strategies has been proposed to address these issues. They differ in at least one of the following three central aspects from traditional AL: (1) They explicitly consider (multiple) human annotators whose performances can be affected by various factors, such as missing expertise. (2) They generalize the interaction with human annotators by considering different query and annotation types, such as asking an annotator for feedback on an inferred classification rule. (3) They take more complex cost schemes regarding annotations and misclassifications into account. This survey provides an overview of these AL strategies and refers to them as real-world AL. Therefore, we introduce a general real-world AL strategy as part of a learning cycle and use its elements, e.g., the query and annotator selection algorithm, to categorize about 60 real-world AL strategies. Finally, we outline possible directions for future research in the field of AL.

📄 PDF Abstract BibTeX arXiv:2109.11301

Code (0)

등록된 구현이 없습니다.

Tasks

Active Learning

Similar Papers 제목 키워드 기반

Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree

2024-10-16 · Harbani Jaggi, Kashyap Murali, Eve Fleisig, Erdem Biyik

When annotators disagree, predicting the labels given by individual annotators can capture nuances overlooked by traditional label aggregation. We introduce three approaches to predicting individual annotator ratings on …

Collaborative FilteringIn-Context LearningPredictionSurvey

Correcting Annotator Bias in Training Data: Population-Aligned Instance Replication (PAIR)

2025-01-12 · Stephanie Eckman, Bolei Ma, Christoph Kern, Rob Chew 외

Models trained on crowdsourced labels may not reflect broader population views, because those who work as annotators do not represent the population. We propose Population-Aligned Instance Replication (PAIR), a method to…

Random Sampling in an Age of Automation: Minimizing Expenditures through Balanced Collection and Annotation

2014-10-26 · Oscar Beijbom

Methods for automated collection and annotation are changing the cost-structures of sampling surveys for a wide range of applications. Digital samples in the form of images or audio recordings can be collected rapidly, a…

MentalQA: An Annotated Arabic Corpus for Questions and Answers of Mental Healthcare

2024-05-21 · Hassan Alhuzali, Ashwag Alasmari, Hamad Alsaleh

Mental health disorders significantly impact people globally, regardless of background, education, or socioeconomic status. However, access to adequate care remains a challenge, particularly for underserved communities w…

AnatomyEpidemiologyQuestion Answering

Towards Assessing Argumentation Annotation - A First Step

2019-08-01 · WS 2019 8 · Anna Lindahl, Lars Borin, Jacobo Rouces

This paper presents a first attempt at using Walton{'}s argumentation schemes for annotating arguments in Swedish political text and assessing the feasibility of using this particular set of schemes with two linguistical…