paper-with-me

홈 › Papers

BLEU, METEOR, BERTScore: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-oriented Text

2021-09-29 · TRITON 2021 7 · Hadeel Saadany, Constantin Orasan

Social media companies as well as authorities make extensive use of artificial intelligence (AI) tools to monitor postings of hate speech, celebrations of violence or profanity. Since AI software requires massive volumes of data to train computers, Machine Translation (MT) of the online content is commonly used to process posts written in several languages and hence augment the data needed for training. However, MT mistakes are a regular occurrence when translating sentiment-oriented user-generated content (UGC), especially when a low-resource language is involved. The adequacy of the whole process relies on the assumption that the evaluation metrics used give a reliable indication of the quality of the translation. In this paper, we assess the ability of automatic quality metrics to detect critical machine translation errors which can cause serious misunderstanding of the affect message. We compare the performance of three canonical metrics on meaningless translations where the semantic content is seriously impaired as compared to meaningful translations with a critical error which exclusively distorts the sentiment of the source text. We conclude that there is a need for fine-tuning of automatic metrics to make them more robust in detecting sentiment critical errors.

📄 PDF Abstract BibTeX arXiv:2109.14250

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Are metrics measuring what they should? An evaluation of image captioning task metrics

2022-07-04 · Othón González-Chávez, Guillermo Ruiz, Daniela Moctezuma, Tania A. Ramirez-delReal

Image Captioning is a current research task to describe the image content using the objects and their relationships in the scene. To tackle this task, two important research areas converge, artificial vision, and natural…

Image Captioning

CTest-Metric: A Unified Framework to Assess Clinical Validity of Metrics for CT Report Generation

2026-01-16 · Vanshali Sharma, Andrea Mia Bejar, Gorkem Durak, Ulas Bagci arxiv

In the generative AI era, where even critical medical tasks are increasingly automated, radiology report generation (RRG) continues to rely on suboptimal metrics for quality assessment. Developing domain-specific metrics…

Generating clickbait spoilers with an ensemble of large language models

2024-05-25 · Mateusz Woźny, Mateusz Lango

Clickbait posts are a widespread problem in the webspace. The generation of spoilers, i.e. short texts that neutralize clickbait by providing information that satisfies the curiosity induced by it, is one of the proposed…

Passage RetrievalQuestion AnsweringRetrieval

Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs

2024-06-04 · Nik Bear Brown

This paper surveys evaluation techniques to enhance the trustworthiness and understanding of Large Language Models (LLMs). As reliance on LLMs grows, ensuring their reliability, fairness, and transparency is crucial. We …

BenchmarkingFairnessFew-Shot LearningHallucination+1

Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case Study

2026-02-09 · Arushi Rai, Adriana Kovashka arxiv

While there is rapid progress in video-LLMs with advanced reasoning capabilities, prior work shows that these models struggle on the challenging task of sports feedback generation and require expensive and difficult-to-c…

Machine TranslationText Generation