paper-with-me

Papers

Improving LLM Leaderboards with Psychometrical Methodology

2025-01-27 · Denis Federiakin

The rapid development of large language models (LLMs) has necessitated the creation of benchmarks to evaluate their performance. These benchmarks resemble human tests and surveys, as they consist of sets of questions designed to measure emergent properties in the cognitive behavior of these systems. However, unlike the well-defined traits and abilities studied in social sciences, the properties measured by these benchmarks are often vaguer and less rigorously defined. The most prominent benchmarks are often grouped into leaderboards for convenience, aggregating performance metrics and enabling comparisons between models. Unfortunately, these leaderboards typically rely on simplistic aggregation methods, such as taking the average score across benchmarks. In this paper, we demonstrate the advantages of applying contemporary psychometric methodologies - originally developed for human tests and surveys - to improve the ranking of large language models on leaderboards. Using data from the Hugging Face Leaderboard as an example, we compare the results of the conventional naive ranking approach with a psychometrically informed ranking. The findings highlight the benefits of adopting psychometric techniques for more robust and meaningful evaluation of LLM performance.

📄 PDF Abstract BibTeX arXiv:2501.17200

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Designing LLM-Agents with Personalities: A Psychometric Approach

2024-10-25 · Muhua Huang, Xijuan Zhang, Christopher Soto, James Evans

This research introduces a novel methodology for assigning quantifiable, controllable and psychometrically validated personalities to Large Language Models-Based Agents (Agents) using the Big Five personality framework. …

Decision Makingvalid

KAQG: A Knowledge-Graph-Enhanced RAG for Difficulty-Controlled Question Generation

2025-05-12 · Ching Han Chen, Ming Fang Shiu

KAQG introduces a decisive breakthrough for Retrieval-Augmented Generation (RAG) by explicitly tackling the two chronic weaknesses of current pipelines: transparent multi-step reasoning and fine-grained cognitive difficu…

Knowledge GraphsQuestion GenerationQuestion-GenerationRAG+2

Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Natural language processing researchers have identified limitations of evaluation methodology for generation tasks, with new questions raised about the validity of automatic metrics and of crowdworker judgments. Meanwhil…

Image CaptioningMachine TranslationText Generation

Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand

2021-12-08 · NAACL 2022 7 · Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan 외

Natural language processing researchers have identified limitations of evaluation methodology for generation tasks, with new questions raised about the validity of automatic metrics and of crowdworker judgments. Meanwhil…

Image CaptioningMachine TranslationText Generation

La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America

2025-07-01 · María Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier 외 arxiv

Leaderboards showcase the current capabilities and limitations of Large Language Models (LLMs). To motivate the development of LLMs that represent the linguistic and cultural diversity of the Spanish-speaking community, …