paper-with-me

Papers

Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation

2026-04-20 · Ran Zhang, Steffen Eger, Arda Tezcan, Wei Zhao, Simone Paolo Ponzetto, Lieve Macken arxiv

Large language models (LLMs) are increasingly used for creative tasks such as literary translation. Yet translational creativity remains underexplored and is rarely evaluated at scale, while source-text comprehension is typically studied in isolation, despite the fact that, in professional translation, comprehension and creativity are tightly intertwined. We address these gaps with a paired-task framework applied to literary excerpts from 11 books. Task 1 assesses source-text comprehension, and Task 2 evaluates translational creativity through Units of Creative Potential (UCPs), such as metaphors and wordplay. Using a scalable evaluation setup that combines expert human annotations with UCP-based automatic scoring, we benchmark 23 models and four creativity-oriented prompts. Our findings show that strong comprehension does not translate into human-level creativity: models often produce literal or contextually inappropriate renderings, with particularly large gaps for the more distant English-Chinese language pair. Creativity-oriented prompts yield only modest gains, and only one model, Mistral-Large, comes close to human-level creativity (0.167 vs. 0.246). Across all model-prompt combinations, only three exceed a creativity score of 0.1, while the rest remain at or near zero.

📄 PDF Abstract BibTeX arXiv:2604.18169

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CyberGym: Evaluating AI Agents' Cybersecurity Capabilities with Real-World Vulnerabilities at Scale

2025-06-03 · Zhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai 외

Large language model (LLM) agents are becoming increasingly skilled at handling cybersecurity tasks autonomously. Thoroughly assessing their cybersecurity capabilities is critical and urgent, given the high stakes in thi…

Large Language Model

Assessing LLM Reliability on Temporally Recent Open-Domain Questions

2026-01-17 · Pushwitha Krishnappa, Amit Das, Vinija Jain, Tathagata Mukherjee 외 arxiv

Large Language Models (LLMs) are increasingly deployed for open-domain question answering, yet their alignment with human perspectives on temporally recent information remains underexplored. We introduce RECOM (Reddit Ev…

Open-Domain Question AnsweringSemantic Similarity

Beyond Calibration: Physically Informed Learning for Raw-to-Raw Mapping

2025-06-10 · Peter Grönquist, Stepan Tulyakov, Dengxin Dai

Achieving consistent color reproduction across multiple cameras is essential for seamless image fusion and Image Processing Pipeline (ISP) compatibility in modern devices, but it is a challenging task due to variations i…

Missing Information, Unresponsive Authors, Experimental Flaws: The Impossibility of Assessing the Reproducibility of Previous Human Evaluations in NLP

2023-05-02 · Anya Belz, Craig Thomson, Ehud Reiter, Gavin Abercrombie 외

We report our efforts in identifying a set of previous human evaluations in NLP that would be suitable for a coordinated study examining what makes human evaluations in NLP more/less reproducible. We present our results …

REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?

2025-07-25 · Chuxuan Hu, Liyun Zhang, Yeji Lim, Aum Wadhwani 외 arxiv

Assessing the reproducibility of social science papers is essential for promoting rigor in research processes, but manual assessment is costly. With recent advances in agentic AI systems (i.e., AI agents), we seek to eva…