On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
This paper investigates the impact of verbose LLM translations on evaluation. We first demonstrate the prevalence of this behavior across several LLM outputs drawn from the WMT 2024 general shared task on machine translation. We then identify the primary triggers of verbosity, including safety, copyright concerns, and insufficient context in short input queries. Finally, we show that ignoring this behavior unfairly penalizes more verbose LLMs according to both automatic and human evaluations, highlighting the need to address this issue for more accurate future evaluations.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
SUMART: SUMmARizing Translation from Wordy to Concise Expression
We propose SUMART, a method for summarizing and compressing the volume of verbose subtitle translations. SUMART is designed for understanding translated captions (e.g., interlingual conversations via subtitle translation…
Large Language ModelTranslationMachine Translation Post-Editing (MTPE) from the Perspective of Translation Trainees: Implications for Translation Pedagogy
This paper introduces data on translation trainees’ perceptions of the MTPE process and implications on training in this field. This study aims to analyse trainees’ performance of three MTPE tasks the English-Polish lang…
Machine TranslationTranslationNMT or SMT: Case Study of a Narrow-domain English-Latvian Post-editing Project
The recent technological shift in machine translation from statistical machine translation (SMT) to neural machine translation (NMT) raises the question of the strengths and weaknesses of NMT. In this paper, we present a…
Machine TranslationNMTTranslationAn Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
With the remarkable success of Vision-Language Models (VLMs) on multimodal tasks, concerns regarding their deployment efficiency have become increasingly prominent. In particular, the number of tokens consumed during the…
Reinforcement LearningText GenerationPrompt Decorators: A Declarative and Composable Syntax for Reasoning, Formatting, and Control in LLMs
Large Language Models (LLMs) are central to reasoning, writing, and decision-support workflows, yet users lack consistent control over how they reason and express outputs. Conventional prompt engineering relies on verbos…
Prompt Engineering