paper-with-me

Papers

Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations

2026-05-13 · Kyo Gerrits, Rik van Noord, Ana Guerberof Arenas arxiv

This article investigates the performance of automatic evaluation metrics (AEMs) and LLM-as-a-judge evaluation on literary translation across multiple languages, genres, and translation modalities. The aim is to assess how well these tools align with professionals when evaluating translation, creativity (creative shifts & errors), and see if they can substitute laborious manual annotations. A dataset of literary translations across three modalities (human translation, machine translation, and post-editing), three genres and three language pairs was created and annotated in detail for creativity by experienced professional literary translators. The results show that both AEMs and LLM-as-a-judge evaluations correlate poorly with professional evaluations on creativity, with LLM-as-a-judge showing a systematic bias in favour of machine-translated texts and penalising creative and culturally appropriate solutions. Moreover, performance is consistently worse for more literary genres such as poetry. This highlights fundamental limitations of current automatic evaluation tools for literary translation and the need to create new tools that do not frequently consider out of routine translations as errors.

📄 PDF Abstract BibTeX arXiv:2605.13596

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

On the stochastics of human and artificial creativity

2024-03-03 · Solve Sæbø, Helge Brovold

What constitutes human creativity, and is it possible for computers to exhibit genuine creativity? We argue that achieving human-level intelligence in computers, or so-called Artificial General Intelligence, necessitates…

Philosophy

Using a CNN Model to Assess Paintings' Creativity

2024-08-02 · Zhehan Zhang, Meihua Qian, Li Luo, Qianyi Gao 외

Assessing artistic creativity has long challenged researchers, with traditional methods proving time-consuming. Recent studies have applied machine learning to evaluate creativity in drawings, but not paintings. Our rese…

model

Creativity and Machine Learning: A Survey

2021-04-06 · Giorgio Franceschelli, Mirco Musolesi

There is a growing interest in the area of machine learning and creativity. This survey presents an overview of the history and the state of the art of computational creativity theories, key machine learning techniques (…

BIG-bench Machine LearningSurvey

Creativity Benchmark: A benchmark for marketing creativity for large language models

2025-09-05 · Ninad Bhat, Kieran Browne, Pip Bingemann arxiv

We introduce Creativity Benchmark, an evaluation framework for large language models (LLMs) in marketing creativity. The benchmark covers 100 brands (12 categories) and three prompt types (Insights, Ideas, Wild Ideas). H…

Co-Operation as an Asymmetric Form of Human-Computer Creativity. Case: Peace Machine

2019-08-01 · WS 2019 8 · Mika H{\"a}m{\"a}l{\"a}inen, Timo Honkela

This theoretical paper identifies a need for a definition of asymmetric co-creativity where creativity is expected from the computational agent but not from the human user. Our co-operative creativity framework takes int…

Form