paper-with-me

홈 › Papers

How not to Lie with a Benchmark: Rearranging NLP Leaderboards

2021-12-02 · NeurIPS Workshop ICBINB 2021 12 · Shavrina Tatiana, Malykh Valentin

Comparison with a human is an essential requirement for a benchmark for it to be a reliable measurement of model capabilities. Nevertheless, the methods for model comparison could have a fundamental flaw - the arithmetic mean of separate metrics is used for all tasks of different complexity, different size of test and training sets. In this paper, we examine popular NLP benchmarks' overall scoring methods and rearrange the models by geometric and harmonic mean (appropriate for averaging rates) according to their reported results. We analyze several popular benchmarks including GLUE, SuperGLUE, XGLUE, and XTREME. The analysis shows that e.g. human level on SuperGLUE is still not reached, and there is still room for improvement for the current models.

📄 PDF Abstract BibTeX arXiv:2112.01342

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

2021-08-01 · ACL 2021 5 · Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor 외

Leaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). R…

Improving LLM Leaderboards with Psychometrical Methodology

2025-01-27 · Denis Federiakin

The rapid development of large language models (LLMs) has necessitated the creation of benchmarks to evaluate their performance. These benchmarks resemble human tests and surveys, as they consist of sets of questions des…

Graph-Transporter: A Graph-based Learning Method for Goal-Conditioned Deformable Object Rearranging Task

2023-02-21 · Yuhong Deng, Chongkun Xia, Xueqian Wang, Lipeng Chen

Rearranging deformable objects is a long-standing challenge in robotic manipulation for the high dimensionality of configuration space and the complex dynamics of deformable objects. We present a novel framework, Graph-T…

Object

Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards

2024-11-07 · Varvara Arzt, Allan Hanbury

This paper investigates the transparency in the creation of benchmarks and the use of leaderboards for measuring progress in NLP, with a focus on the relation extraction (RE) task. Existing RE benchmarks often suffer fro…

RelationRelation Extraction

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

2026-05-27 · Pouya Sadeghi, Anamaria Crisan, Jimmy Lin arxiv

Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite known limitations in their reliability and robustness. Yet how they sha…