paper-with-me

홈 › Papers

The Glass Ceiling of Automatic Evaluation in Natural Language Generation

2022-08-31 · Pierre Colombo, Maxime Peyrard, Nathan Noiry, Robert West, Pablo Piantanida

Automatic evaluation metrics capable of replacing human judgments are critical to allowing fast development of new methods. Thus, numerous research efforts have focused on crafting such metrics. In this work, we take a step back and analyze recent progress by comparing the body of existing automatic metrics and human metrics altogether. As metrics are used based on how they rank systems, we compare metrics in the space of system rankings. Our extensive statistical analysis reveals surprising findings: automatic metrics -- old and new -- are much more similar to each other than to humans. Automatic metrics are not complementary and rank systems similarly. Strikingly, human metrics predict each other much better than the combination of all automatic metrics used to predict a human metric. It is surprising because human metrics are often designed to be independent, to capture different aspects of quality, e.g. content fidelity or readability. We provide a discussion of these findings and recommendations for future work in the field of evaluation.

📄 PDF Abstract BibTeX arXiv:2208.14585

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

The glass ceiling in NLP

2018-10-01 · EMNLP 2018 10 · Natalie Schluter

In this paper, we provide empirical evidence based on a rigourously studied mathematical model for bi-populated networks, that a glass ceiling within the field of NLP has developed since the mid 2000s.

Named Entity Recognition -- Is there a glass ceiling?

2019-10-06 · Tomasz Stanislawek, Anna Wróblewska, Alicja Wójcicka, Daniel Ziembicki 외

Recent developments in Named Entity Recognition (NER) have resulted in better and better models. However, is there a glass ceiling? Do we know which types of errors are still hard or even impossible to correct? In this p…

Diagnosticnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Named Entity Recognition - Is There a Glass Ceiling?

2019-11-01 · CONLL 2019 11 · Tomasz Stanislawek, Anna Wr{\'o}blewska, Alicja W{\'o}jcicka, Daniel Ziembicki 외

Recent developments in Named Entity Recognition (NER) have resulted in better and better models. However, is there a glass ceiling? Do we know which types of errors are still hard or even impossible to correct? In this p…

BIG-bench Machine Learningnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Breaking Through the 80\% Glass Ceiling: Raising the State of the Art in Word Sense Disambiguation by Incorporating Knowledge Graph Information

2020-07-01 · ACL 2020 6 · Michele Bevilacqua, Roberto Navigli

Neural architectures are the current state of the art in Word Sense Disambiguation (WSD). However, they make limited use of the vast amount of relational information encoded in Lexical Knowledge Bases (LKB). We present E…

AllWord Sense Disambiguation

Attenuation of Several Common Building Materials in Millimeter-Wave Frequency Bands: 28, 73 and 91 GHz

2020-04-27 · Nozhan Hosseini, Mahfuza Khatun, Changyu Guo, Kairui Du 외

Future cellular systems will make use of millimeter wave (mmWave) frequency bands. Many users in these bands are located indoors, i.e., inside buildings, homes, and offices. Typical building material attenuations in thes…