paper-with-me

홈 › Papers

Perceived Confidence Scoring for Data Annotation with Zero-Shot LLMs

2025-02-11 · Sina Salimian, Gias Uddin, Most Husne Jahan, Shaina Raza

Zero-shot LLMs are now also used for textual classification tasks, e.g., sentiment/emotion detection of a given input as a sentence/article. However, their performance can be suboptimal in such data annotation tasks. We introduce a novel technique Perceived Confidence Scoring (PCS) that evaluates LLM's confidence for its classification of an input by leveraging Metamorphic Relations (MRs). The MRs generate semantically equivalent yet textually mutated versions of the input. Following the principles of Metamorphic Testing (MT), the mutated versions are expected to have annotation labels similar to the input. By analyzing the consistency of LLM responses across these variations, PCS computes a confidence score based on the frequency of predicted labels. PCS can be used both for single LLM and multiple LLM settings (e.g., majority voting). We introduce an algorithm Perceived Differential Evolution (PDE) that determines the optimal weights assigned to the MRs and the LLMs for a classification task. Empirical evaluation shows PCS significantly improves zero-shot accuracy for Llama-3-8B-Instruct (4.96%) and Mistral-7B-Instruct-v0.3 (10.52%), with Gemma-2-9b-it showing a 9.39% gain. When combining all three models, PCS significantly outperforms majority voting by 7.75%.

📄 PDF Abstract BibTeX arXiv:2502.07186

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Emotion Ratings: How Intensity, Annotation Confidence and Agreements are Entangled

2021-03-02 · EACL (WASSA) 2021 4 · Enrica Troiano, Sebastian Padó, Roman Klinger

When humans judge the affective content of texts, they also implicitly assess the correctness of such judgment, that is, their confidence. We hypothesize that people's (in)confidence that they performed well in an annota…

Diagnostic

Clearing noisy annotations for computed tomography imaging

2018-07-23 · Roman Khudorozhkov, Alexander Koryagin, Alexey Kozhevin

One of the problems on the way to successful implementation of neural networks is the quality of annotation. For instance, different annotators can annotate images in a different way and very often their decisions do not…

Computed Tomography (CT)SegmentationSemantic Segmentation

Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization

2025-10-20 · Yuanli Wu, Long Zhang, Yue Du, Bin Li arxiv

We propose a rubric-guided, pseudo-labeled, and prompt-driven zero-shot video summarization framework that bridges large language models with structured semantic reasoning. A small subset of human annotations is converte…

Video Summarization

Preventing Critical Scoring Errors in Short Answer Scoring with Confidence Estimation

2020-07-01 · ACL 2020 6 · Hiroaki Funayama, Shota Sasaki, Yuichiroh Matsubayashi, Tomoya Mizumoto 외

Many recent Short Answer Scoring (SAS) systems have employed Quadratic Weighted Kappa (QWK) as the evaluation measure of their systems. However, we hypothesize that QWK is unsatisfactory for the evaluation of the SAS sys…

CLIPScope: Enhancing Zero-Shot OOD Detection with Bayesian Scoring

2024-05-23 · Hao Fu, Naman Patel, Prashanth Krishnamurthy, Farshad Khorrami

Detection of out-of-distribution (OOD) samples is crucial for safe real-world deployment of machine learning models. Recent advances in vision language foundation models have made them capable of detecting OOD samples wi…