paper-with-me

Papers

When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content

2025-12-19 · Lydia Nishimwe, Benoît Sagot, Rachel Bawden arxiv

User-generated content (UGC) is characterised by frequent use of non-standard language, from spelling errors to expressive choices such as slang, character repetitions, and emojis. This makes evaluating UGC translation challenging: what counts as a "good" translation depends on the desired standardness level of the output. To explore this, we examine the human translation guidelines of four UGC datasets, and derive a taxonomy of twelve non-standard phenomena and five translation actions (NORMALISE, COPY, TRANSFER, OMIT, CENSOR). Our analysis reveals notable differences in how UGC is treated, resulting in a spectrum of standardness in reference translations. We show that translation scores of large language models are highly sensitive to prompts with explicit UGC translation instructions, and that they improve when they align with the dataset guidelines. We argue that fair evaluation requires both models and metrics to be aware of translation guidelines. Finally, we call for clear guidelines during dataset creation and for the development of controllable, guideline-aware evaluation frameworks for UGC translation.

📄 PDF Abstract BibTeX arXiv:2512.17738

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Gold Standard Dependency Corpus for English

2014-05-01 · LREC 2014 5 · Natalia Silveira, Timothy Dozat, Marie-Catherine de Marneffe, Samuel Bowman 외

We present a gold standard annotation of syntactic dependencies in the English Web Treebank corpus using the Stanford Dependencies formalism. This resource addresses the lack of a gold standard dependency treebank for En…

Sentiment Analysis

Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap

2026-07-16 · Olivier Jeunen arxiv

Online controlled experiments are the gold standard for hypothesis testing in online platforms. Notwithstanding their ubiquity, they are notoriously expensive to run, and issues of variance hamper statistical power in as…

Information Retrieval

Multilingual Controlled Generation And Gold-Standard-Agnostic Evaluation of Code-Mixed Sentences

2024-10-14 · Ayushman Gupta, Akhil Bhogal, Kripabandhu Ghosh

Code-mixing, the practice of alternating between two or more languages in an utterance, is a common phenomenon in multilingual communities. Due to the colloquial nature of code-mixing, there is no singular correct way to…

SentenceText Generation

Centroids: Gold standards with distributional variation

2012-05-01 · LREC 2012 5 · Ian Lewin, {\c{S}}enay Kafkas, Dietrich Rebholz-Schuhmann

Motivation: Gold Standards for named entities are, ironically, not standard themselves. Some specify the “one perfect annotation”. Others specify “perfectly good alternatives”. The concept of Silver standard is relativel…

Named Entity Recognition (NER)

Persian Abstract Meaning Representation: Annotation Guidelines and Gold Standard Dataset

2022-05-16 · Reza Takhshid, Tara Azin, Razieh Shojaei, Mohammad bahrani

This paper introduces the Persian Abstract Meaning Representation (AMR) guidelines, a detailed guide for annotating Persian sentences with AMR, focusing on the necessary adaptations to fit Persian's unique syntactic stru…

Abstract Meaning RepresentationSentenceTranslation