Towards A Reliable Ground-Truth For Biased Language Detection
Reference texts such as encyclopedias and news articles can manifest biased language when objective reporting is substituted by subjective writing. Existing methods to detect bias mostly rely on annotated data to train machine learning models. However, low annotator agreement and comparability is a substantial drawback in available media bias corpora. To evaluate data collection options, we collect and compare labels obtained from two popular crowdsourcing platforms. Our results demonstrate the existing crowdsourcing approaches' lack of data quality, underlining the need for a trained expert framework to gather a more reliable dataset. By creating such a framework and gathering a first dataset, we are able to improve Krippendorff's $\alpha$ = 0.144 (crowdsourcing labels) to $\alpha$ = 0.419 (expert labels). We conclude that detailed annotator training increases data quality, improving the performance of existing bias detection systems. We will continue to extend our dataset in the future.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesBias DetectionSimilar Papers 제목 키워드 기반
Unbiased IoU for Spherical Image Object Detection
As one of the most fundamental and challenging problems in computer vision, object detection tries to locate object instances and find their categories in natural images. The most important step in the evaluation of obje…
Objectobject-detectionObject DetectionBest Arm Identification with LLM Judges and Limited Human
We study fixed-confidence best-arm identification (BAI) where a cheap but potentially biased proxy (e.g., LLM judge) is available for every sample, while an expensive ground-truth label can only be acquired selectively w…
3DRef: 3D Dataset and Benchmark for Reflection Detection in RGB and Lidar Data
Reflective surfaces present a persistent challenge for reliable 3D mapping and perception in robotics and autonomous systems. However, existing reflection datasets and benchmarks remain limited to sparse 2D data. This pa…
Image SegmentationMirror DetectionPoint Cloud SegmentationSegmentation+1Self-supervised conformal prediction for uncertainty quantification in Poisson imaging problems
Image restoration problems are often ill-posed, leading to significant uncertainty in reconstructed images. Accurately quantifying this uncertainty is essential for the reliable interpretation of reconstructed images. Ho…
Conformal PredictionDeblurringDenoisingImage Denoising+3Don't Throw it Away! The Utility of Unlabeled Data in Fair Decision Making
Decision making algorithms, in practice, are often trained on data that exhibits a variety of biases. Decision-makers often aim to take decisions based on some ground-truth target that is assumed or expected to be unbias…
Decision MakingFairness