paper-with-me

Papers

Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry

2025-10-17 · Bolei Ma, Yina Yao, Anna-Carolina Haensch arxiv

Large Language Models (LLMs) are increasingly applied to creative domains, yet their performance in classical Chinese poetry generation and evaluation remains poorly understood. We propose a three-step evaluation framework that combines computational metrics, LLM-as-a-judge assessment, and human expert validation. Using this framework, we evaluate six state-of-the-art LLMs across multiple dimensions of poetic quality, including themes, emotions, imagery, form, and style, in the context of Tang poetry generation. Our analysis reveals a critical "echo chamber" effect: LLMs systematically overrate machine-generated poems that mimic statistical patterns yet fail strict prosodic rules, diverging significantly from human expert judgments. These findings underscore the limitations of using LLMs as standalone evaluators for culturally complex tasks, highlighting the necessity of hybrid human-model validation frameworks.

📄 PDF Abstract BibTeX arXiv:2510.15313

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Behavioral Bias of Vision-Language Models: A Behavioral Finance View

2024-09-23 · Yuhang Xiao, Yudi Lin, Ming-Chang Chiu

Large Vision-Language Models (LVLMs) evolve rapidly as Large Language Models (LLMs) was equipped with vision modules to create more human-like models. However, we should carefully evaluate their applications in different…

An Empirical Analysis on Large Language Models in Debate Evaluation

2024-05-28 · Xinyi Liu, Pinxin Liu, Hangfeng He

In this study, we investigate the capabilities and inherent biases of advanced large language models (LLMs) such as GPT-3.5 and GPT-4 in the context of debate evaluation. We discover that LLM's performance exceeds humans…

ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models

2024-05-28 · Aparna Elangovan, Ling Liu, Lei Xu, Sravan Bodapati 외

In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon insights from disciplines such as user experience research and h…

Experimental Design

DateLogicQA: Benchmarking Temporal Biases in Large Language Models

2024-12-17 · Gagan Bhatia, MingZe Tang, Cristina Mahanta, Madiha Kazi

This paper introduces DateLogicQA, a benchmark with 190 questions covering diverse date formats, temporal contexts, and reasoning types. We propose the Semantic Integrity Metric to assess tokenization quality and analyse…

Benchmarking

PerQ: Efficient Evaluation of Multilingual Text Personalization Quality

2025-09-30 · Dominik Macko, Andrew Pulver arxiv

Since no metrics are available to evaluate specific aspects of a text, such as its personalization quality, the researchers often rely solely on large language models to meta-evaluate such texts. Due to internal biases o…