paper-with-me

Papers

Can LLMs Rank? A Tale of Triads and Triage

2026-06-29 · Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das arxiv

From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources. Ranking large groups simultaneously is cognitively demanding and error-prone. A natural solution, drawing on decades of social choice theory, elicits pairwise comparisons and aggregates them into a total order. However, a fundamental question remains when LLMs serve as the pairwise judge: how can a practitioner tell, before committing to a ranking, whether the LLM's judgments are sufficiently consistent to trust the result? We discuss two different ways of identifying consistency. A classical diagnostic, the coefficient of consistency $ζ$, originally developed to measure judge reliability by counting circular triads in tournament graphs, provides a cheap, model-free measure of intra-run consistency. Various standard measures of distance between rankings, for example Kendall's $τ$, can measure inter-run variability. We show, in both theory and practice, that these measures are independently valuable, and advocate for using both to assess reliability of rankings. We demonstrate the practical importance of our results across two high-stakes prioritization tasks: homelessness service allocation and emergency department triage. Three different leading LLMs have considerably different performance profiles across these two axes of consistency. We provide guidelines for how practitioners could think about measuring and assessing consistency before committing to a model for ranking or prioritization.

📄 PDF Abstract BibTeX arXiv:2606.30412

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Axiomatizations of inconsistency indices for triads

2018-01-10 · László Csató

Pairwise comparison matrices often exhibit inconsistency, therefore many indices have been suggested to measure their deviation from a consistent matrix. A set of axioms has been proposed recently that is required to be …

TriageRA-CCF: Source-Side Clinical Confidence and Coverage Signals for Adaptive Rank Budgeting in Medical LLMs

2026-06-28 · Shucan Ji, Yining Huang, Hongliang Guo arxiv

Medical large language models are commonly adapted with a fixed low-rank budget, even though medical questions differ substantially in confidence, clinical coverage, and cross-domain difficulty. We study adaptive rank bu…

Question Answering

Medical Triage as Pairwise Ranking: A Benchmark for Urgency in Patient Portal Messages

2026-01-19 · Joseph Gatto, Parker Seegmiller, Timothy Burdick, Philip Resnik 외 arxiv

Medical triage is the task of allocating medical resources and prioritizing patients based on medical need. This paper introduces the first large-scale public dataset for studying medical triage in the context of asynchr…

LLM-RankFusion: Mitigating Intrinsic Inconsistency in LLM-based Ranking

2024-05-31 · Yifan Zeng, Ojas Tendolkar, Raymond Baartmans, Qingyun Wu 외

Ranking passages by prompting a large language model (LLM) can achieve promising performance in modern information retrieval (IR) systems. A common approach to sort the ranking list is by prompting LLMs for a pairwise or…

In-Context LearningInformation RetrievalLanguage ModellingLarge Language Model

TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings

2025-08-23 · Md Afif Al Mamun, Gias Uddin, Lan Xia, Longyu Zhang arxiv

Pretrained Language Models or PLMs are transformer-based architectures that can be used in bug triaging tasks. PLMs can better capture token semantics than traditional Machine Learning (ML) models that rely on statistica…