paper-with-me

홈 › Papers

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance

2024-10-24 · Omer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor, Roi Reichart

NLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field. Traditionally, expert annotations ensure high-quality labels; however, the cost of expert annotation does not scale well with the growing demand for larger datasets required by modern models. While crowd-sourcing provides a more scalable solution, it often comes at the expense of annotation precision and consistency. Recent advancements in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets. In this work, we consider the recent approach of LLM-as-a-judge, leveraging an ensemble of LLMs to flag potentially mislabeled examples. Through a case study of four datasets from the TRUE benchmark, covering different tasks and domains, we empirically analyze the labeling quality of existing datasets, and compare expert, crowd-sourced, and our LLM-based annotations in terms of agreement, label quality, and efficiency, demonstrating the strengths and limitations of each annotation method. Our findings reveal a substantial number of label errors, which, when corrected, induce a significant upward shift in reported model performance. This suggests that many of the LLMs so-called mistakes are due to label errors rather than genuine model failures. Additionally, we discuss the implications of mislabeled data and propose methods to mitigate them in training to improve model performance.

📄 PDF Abstract BibTeX arXiv:2410.18889

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interpreting Chest X-rays via CNNs that Exploit Hierarchical Disease Dependencies and Uncertainty Labels

2020-05-25 · MIDL 2019 7 · Hieu H. Pham, Tung T. Le, Dat T. Ngo, Dat Q. Tran 외

The chest X-rays (CXRs) is one of the views most commonly ordered by radiologists (NHS),which is critical for diagnosis of many different thoracic diseases. Accurately detecting thepresence of multiple diseases from CXRs…

Generalist Segmentation Algorithm for Photoreceptors Analysis in Adaptive Optics Imaging

2024-08-27 · Mikhail Kulyabin, Aline Sindel, Hilde Pedersen, Stuart Gilson 외

Analyzing the cone photoreceptor pattern in images obtained from the living human retina using quantitative methods can be crucial for the early detection and management of various eye conditions. Confocal adaptive optic…

MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models

2024-09-29 · Vibhor Agarwal, Yiqiao Jin, Mohit Chandra, Munmun De Choudhury 외

The remarkable capabilities of large language models (LLMs) in language understanding and generation have not rendered them immune to hallucinations. LLMs can still generate plausible-sounding but factually incorrect or …

Hallucination

Understanding Fine-grained Distortions in Reports of Scientific Findings

2024-02-19 · Amelie Wührl, Dustin Wright, Roman Klinger, Isabelle Augenstein

Distorted science communication harms individuals and society as it can lead to unhealthy behavior change and decrease trust in scientific institutions. Given the rapidly increasing volume of science communication in rec…

Articles

Interpreting chest X-rays via CNNs that exploit hierarchical disease dependencies and uncertainty labels

2019-11-15 · Hieu H. Pham, Tung T. Le, Dat Q. Tran, Dat T. Ngo 외

Chest radiography is one of the most common types of diagnostic radiology exams, which is critical for screening and diagnosis of many different thoracic diseases. Specialized algorithms have been developed to detect sev…

DiagnosticMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION