paper-with-me

홈 › Papers

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

2026-05-27 · Pouya Sadeghi, Anamaria Crisan, Jimmy Lin arxiv

Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite known limitations in their reliability and robustness. Yet how they shape researchers' actual practice remains empirically uncharted. We address this gap through semi-structured interviews with eight researchers across four computer science subfields, analyzed using reflexive thematic analysis. We find a near-universal paradox of pragmatic skepticism: while participants expressed deep distrust of leaderboard rankings, they continued to use them as rough decision-making aids. Peer networks, not leaderboards, emerged as the primary model selection mechanism, and arena-based (human-voting) leaderboards were consistently preferred over static benchmark leaderboards. Leaderboard influence varied sharply across subfields, revealing that disciplinary culture, not individual attitudes, mediates engagement; for instance, NLP researchers faced state-of-the-art comparison pressure while HCI and Systems/Privacy researchers reported none. Across these differences, however, participants converged on cost transparency as the most demanded missing feature (seven of eight). We translate these findings into concrete design recommendations that align evaluation infrastructure with how researchers actually use it, such as task-specific score breakdowns, cost integration, and voter-demographic disclosure.

📄 PDF Abstract BibTeX arXiv:2605.28966

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The AI Alignment Paradox

2024-05-31 · Robert West, Roland Aydin

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today'…

AI Should Be More Human, Not More Complex

2025-07-27 · Carlo Esposito arxiv

Large Language Models (LLMs) in search applications increasingly prioritize verbose, lexically complex responses that paradoxically reduce user satisfaction and engagement. Through a comprehensive study of 10.000 (est.) …

ORKG-Leaderboards: A Systematic Workflow for Mining Leaderboards as a Knowledge Graph

2023-05-10 · Salomon Kabongo, Jennifer D'Souza, Sören Auer

The purpose of this work is to describe the Orkg-Leaderboard software designed to extract leaderboards defined as Task-Dataset-Metric tuples automatically from large collections of empirical research papers in Artificial…

On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards

2024-07-04 · Zhimin Zhao, Abdul Ali Bangash, Filipe Roseiro Côgo, Bram Adams 외

Foundation models (FM), such as large language models (LLMs), which are large-scale machine learning (ML) models, have demonstrated remarkable adaptability in various downstream software engineering (SE) tasks, such as c…

Code Completion

The Psychology of Learning from Machines: Anthropomorphic AI and the Paradox of Automation in Education

2026-01-07 · Junaid Qadir, Muhammad Mumtaz arxiv

As AI tutors enter classrooms at unprecedented speed, their deployment increasingly outpaces our grasp of the psychological and social consequences of such technology. Yet decades of research in automation psychology, hu…