Modeling Human Annotation Errors to Design Bias-Aware Systems for Social Stream Processing
High-quality human annotations are necessary to create effective machine learning systems for social media. Low-quality human annotations indirectly contribute to the creation of inaccurate or biased learning systems. We show that human annotation quality is dependent on the ordering of instances shown to annotators (referred as 'annotation schedule'), and can be improved by local changes in the instance ordering provided to the annotators, yielding a more accurate annotation of the data stream for efficient real-time social media analytics. We propose an error-mitigating active learning algorithm that is robust with respect to some cases of human errors when deciding an annotation schedule. We validate the human error model and evaluate the proposed algorithm against strong baselines by experimenting on classification tasks of relevant social media posts during crises. According to these experiments, considering the order in which data instances are presented to human annotators leads to both an increase in accuracy for machine learning and awareness toward some potential biases in human learning that may affect the automated classifier.
Code (0)
등록된 구현이 없습니다.
Tasks
Active LearningBIG-bench Machine LearningSimilar Papers 제목 키워드 기반
Modeling Annotator Preference and Stochastic Annotation Error for Medical Image Segmentation
Manual annotation of medical images is highly subjective, leading to inevitable and huge annotation biases. Deep learning models may surpass human performance on a variety of tasks, but they may also mimic or amplify the…
Image SegmentationMedical Image SegmentationSegmentationSemantic SegmentationSurrogate Representation Inference for Text and Image Annotations
As researchers increasingly rely on machine learning models and LLMs to annotate unstructured data, such as texts or images, various approaches have been proposed to correct bias in downstream statistical analysis. Howev…
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected from human annotators under controlled …
Reinforcement LearningEvaluating and Correcting Human Annotation Bias in Dynamic Micro-Expression Recognition
Existing manual labeling of micro-expressions is subject to errors in accuracy, especially in cross-cultural scenarios where deviation in labeling of key frames is more prominent. To address this issue, this paper presen…
Micro-Expression RecognitionBiasLab: Toward Explainable Political Bias Detection with Dual-Axis Annotations and Rationale Indicators
We present BiasLab, a dataset of 300 political news articles annotated for perceived ideological bias. These articles were selected from a curated 900-document pool covering diverse political events and source biases. Ea…
ArticlesBias Detection