paper-with-me

Papers

Language Agnostic Automatic Summarization Evaluation

2020-05-01 · LREC 2020 5 · Christopher Tauchmann, Margot Mieskes

So far work on automatic summarization has dealt primarily with English data. Accordingly, evaluation methods were primarily developed with this language in mind. In our work, we present experiments of adapting available evaluation methods such as ROUGE and PYRAMID to non-English data. We base our experiments on various English and non-English homogeneous benchmark data sets as well as a non-English heterogeneous data set. Our results indicate that ROUGE can indeed be adapted to non-English data {--} both homogeneous and heterogeneous. Using a recent implementation of performing an automatic PYRAMID evaluation, we also show its adaptability to non-English data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation

2022-12-15 · Yixin Liu, Alexander R. Fabbri, PengFei Liu, Yilun Zhao 외

Human evaluation is the foundation upon which the evaluation of both summarization systems and automatic metrics rests. However, existing human evaluation studies for summarization either exhibit a low inter-annotator ag…

NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist

2023-05-15 · Iftitahu Ni'mah, Meng Fang, Vlado Menkovski, Mykola Pechenizkiy

In this study, we analyze automatic evaluation metrics for Natural Language Generation (NLG), specifically task-agnostic metrics and human-aligned metrics. Task-agnostic metrics, such as Perplexity, BLEU, BERTScore, are …

Controllable Language ModellingDialogue GenerationLanguage Modellingnlg evaluation+3

Revisiting Automatic Question Summarization Evaluation in the Biomedical Domain

2023-03-18 · Hongyi Yuan, Yaoyun Zhang, Fei Huang, Songfang Huang

Automatic evaluation metrics have been facilitating the rapid development of automatic summarization methods by providing instant and fair assessments of the quality of summaries. Most metrics have been developed for the…

Text Generation

A Comparative Study of Quality Evaluation Methods for Text Summarization

2024-06-30 · Huyen Nguyen, Haihua Chen, Lavanya Pobbathi, Junhua Ding

Evaluating text summarization has been a challenging task in natural language processing (NLP). Automatic metrics which heavily rely on reference summaries are not suitable in many situations, while human evaluation is t…

Text Summarization

Towards Human-Free Automatic Quality Evaluation of German Summarization

2021-05-13 · Neslihan Iskender, Oleg Vasilyev, Tim Polzehl, John Bohannon 외

Evaluating large summarization corpora using humans has proven to be expensive from both the organizational and the financial perspective. Therefore, many automatic evaluation metrics have been developed to measure the s…

InformativenessLanguage ModelingLanguage Modelling