paper-with-me

홈 › Papers

A Framework for Evaluating LLMs Under Task Indeterminacy

2024-11-21 · Luke Guerdan, Hanna Wallach, Solon Barocas, Alexandra Chouldechova

Large language model (LLM) evaluations often assume there is a single correct response -- a gold label -- for each item in the evaluation corpus. However, some tasks can be ambiguous -- i.e., they provide insufficient information to identify a unique interpretation -- or vague -- i.e., they do not clearly indicate where to draw the line when making a determination. Both ambiguity and vagueness can cause task indeterminacy -- the condition where some items in the evaluation corpus have more than one correct response. In this paper, we develop a framework for evaluating LLMs under task indeterminacy. Our framework disentangles the relationships between task specification, human ratings, and LLM responses in the LLM evaluation pipeline. Using our framework, we conduct a synthetic experiment showing that evaluations that use the "gold label" assumption underestimate the true performance. We also provide a method for estimating an error-adjusted performance interval given partial knowledge about indeterminate items in the evaluation corpus. We conclude by outlining implications of our work for the research community.

📄 PDF Abstract BibTeX arXiv:2411.13760

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Evaluating Information Loss in Temporal Dependency Trees

2020-05-01 · LREC 2020 5 · Mustafa Ocal, Mark Finlayson

Temporal Dependency Trees (TDTs) have emerged as an alternative to full temporal graphs for representing the temporal structure of texts, with a key advantage being that TDTs can be straightforwardly computed using adapt…

Preference Reasoning under Indeterminacy in Large Language Models

2026-08-19 · Hadi Hosseini, Samarth Khanna, Xiyuan Wang arxiv

As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world…

The Fuzzy ROC

2019-03-04 · Giovanni Parmigiani

The fuzzy ROC extends Receiver Operating Curve (ROC) visualization to the situation where some data points, falling in an indeterminacy region, are not classified. It addresses two challenges: definition of sensitivity a…

SensitivitySpecificity

The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation

2020-07-01 · Yuma Koizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada 외

This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captioning. Our submission focuses on solving two…

Audio captioningCaption GenerationMulti-Task LearningSentence

Stochastic Trajectory Prediction via Motion Indeterminacy Diffusion

2022-03-25 · CVPR 2022 1 · Tianpei Gu, Guangyi Chen, Junlong Li, Chunze Lin 외

Human behavior has the nature of indeterminacy, which requires the pedestrian trajectory prediction system to model the multi-modality of future motion states. Unlike existing stochastic trajectory prediction methods whi…

DiversityPedestrian Trajectory PredictionPredictionTrajectory Prediction