paper-with-me

홈 › Papers

Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation

2024-07-03 · Paulo Cavalin, Pedro Henrique Domingues, Claudio Pinhanez

In this paper we show that corpus-level aggregation hinders considerably the capability of lexical metrics to accurately evaluate machine translation (MT) systems. With empirical experiments we demonstrate that averaging individual segment-level scores can make metrics such as BLEU and chrF correlate much stronger with human judgements and make them behave considerably more similar to neural metrics such as COMET and BLEURT. We show that this difference exists because corpus- and segment-level aggregation differs considerably owing to the classical average of ratio versus ratio of averages Mathematical problem. Moreover, as we also show, such difference affects considerably the statistical robustness of corpus-level aggregation. Considering that neural metrics currently only cover a small set of sufficiently-resourced languages, the results in this paper can help make the evaluation of MT systems for low-resource languages more trustworthy.

📄 PDF Abstract BibTeX arXiv:2407.12832

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentence

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Acoustic Analysis of Native (L1) Bengali Speakers’ Phonological Realization of English Lexical Stress Contrast

2020-12-01 · ICON 2020 12 · Shambhu Nath Saha, Shyamal Kr. Das Mandal

Acoustically, English lexical stress is multidimensional and involving manipulation of duration, intensity, fundamental frequency (F0) and vowel quality. The current study investigates the acquisition of English lexical …

Sentence

Controllable Text Simplification with Lexical Constraint Loss

2019-07-01 · ACL 2019 7 · Daiki Nishihara, Tomoyuki Kajiwara, Yuki Arase

We propose a method to control the level of a sentence in a text simplification task. Text simplification is a monolingual translation task translating a complex sentence into a simpler and easier to understand the alter…

SentenceText SimplificationTranslation

SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

2025-11-21 · Shrikant Kendre, Austin Xu, Honglu Zhou, Michael Ryoo 외 arxiv

Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed…

Visual Question Answering

Learning-based Composite Metrics for Improved Caption Evaluation

2018-07-01 · ACL 2018 7 · Naeha Sharif, Lyndon White, Mohammed Bennamoun, Syed Afaq Ali Shah

The evaluation of image caption quality is a challenging task, which requires the assessment of two main aspects in a caption: adequacy and fluency. These quality aspects can be judged using a combination of several ling…

Image CaptioningLanguage ModelingLanguage ModellingSemantic Textual Similarity+1

BLEU is Not Suitable for the Evaluation of Text Simplification

2018-10-14 · EMNLP 2018 10 · Elior Sulem, Omri Abend, Ari Rappoport

BLEU is widely considered to be an informative metric for text-to-text generation, including Text Simplification (TS). TS includes both lexical and structural aspects. In this paper we show that BLEU is not suitable for …

SentenceText GenerationText Simplification