paper-with-me

Papers

Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory

2025-05-21 · Hongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Wei Bao, Jian Dong, Bing Xu, Conghui Zhu, Hailong Cao, Tiejun Zhao

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining main-stream prominent LLM benchmarks using results from diverse models. We first propose a new framework for accurate and reliable estimations of item characteristics and model abilities. Specifically, we propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. Based on PSN-IRT, we conduct extensive analysis which reveals significant and varied shortcomings in the measurement quality of current benchmarks. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference.

📄 PDF Abstract BibTeX arXiv:2505.15055

Code (1)

Joe-Hall-Lee/PSN-IRT 공식 구현 pytorch

Tasks

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking

2026-03-20 · Tomas Ruiz, Tanalp Agustoslu, Carsten Schwemmer arxiv

Human Label Variation (HLV), i.e. systematic differences among annotators' judgments, remains underexplored in benchmarks despite rapid progress in large language model (LLM) development. We address this gap by introduci…

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification

2026-04-09 · Xun Zhu, Fanbin Mo, Xi Chen, Kaili Zheng 외 arxiv

The rise of multimodal large language models (MLLMs) has sparked an unprecedented wave of applications in the field of medical imaging analysis. However, as one of the earliest and most fundamental tasks integrated into …

Medical Image Classification

When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs

2026-02-03 · Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev 외 arxiv

Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning. However, objective evaluation of retrosynthesis performance remains limited. Existing benchmarks and m…

Single-step retrosynthesisDrug Discovery

Benchmarking that Matters: Rethinking Benchmarking for Practical Impact

2025-11-15 · Anna V. Kononova, Niki van Stein, Olaf Mersmann, Thomas Bäck 외 arxiv

Benchmarking has driven scientific progress in Evolutionary Computation, yet current practices fall short of real-world needs. Widely used synthetic suites such as BBOB and CEC isolate algorithmic phenomena but poorly re…

Dynabench: Rethinking Benchmarking in NLP

2021-04-07 · NAACL 2021 4 · Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik 외

We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking. Dynabench runs in a web browser and supports human-and-model-in-the-loop dataset creation: annotators seek to create ex…

Benchmarking