paper-with-me

홈 › Papers

The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario

2024-06-22 · Carlos Gómez-Rodríguez, Paul Williams

This is a summary of the paper "A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing", which was published in Findings of EMNLP 2023. We evaluate a range of recent state-of-the-art, instruction-tuned large language models (LLMs) on an English creative writing task, and compare them to human writers. For this purpose, we use a specifically-tailored prompt (based on an epic combat between Ignatius J. Reilly, main character of John Kennedy Toole's "A Confederacy of Dunces", and a pterodactyl) to minimize the risk of training data leakage and force the models to be creative rather than reusing existing stories. The same prompt is presented to LLMs and human writers, and evaluation is performed by humans using a detailed rubric including various aspects like fluency, style, originality or humor. Results show that some state-of-the-art commercial LLMs match or slightly outperform our human writers in most of the evaluated dimensions. Open-source LLMs lag behind. Humans keep a close lead in originality, and only the top three LLMs can handle humor at human-like levels.

📄 PDF Abstract BibTeX arXiv:2406.15891

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing?

2024-07-01 · Guillermo Marco, Julio Gonzalo, Ramón del Castillo, María Teresa Mateo Girona

It has become routine to report research results where Large Language Models (LLMs) outperform average humans in a wide range of language-related tasks, and creative text writing is no exception. It seems natural, then, …

Art or Artifice? Large Language Models and the False Promise of Creativity

2023-09-25 · Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan 외

Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by …

Evaluating Creative Short Story Generation in Humans and Large Language Models

2024-11-04 · Mete Ismayilzada, Claire Stevenson, Lonneke van der Plas

Story-writing is a fundamental aspect of human imagination, relying heavily on creativity to produce narratives that are novel, effective, and surprising. While large language models (LLMs) have demonstrated the ability …

DiversitySentenceStory Generation

Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback

2025-07-21 · Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella Lapata arxiv

Can LLMs provide support to creative writers by giving meaningful writing feedback? In this paper, we explore the challenges and limitations of model-generated writing feedback by defining a new task, dataset, and evalua…

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

2025-07-01 · Daniel Fein, Sebastian Russo, Violet Xiang, Kabir Jolly 외 arxiv

Evaluating creative writing generated by large language models (LLMs) remains challenging because open-ended narratives lack ground truths. Without performant automated evaluation methods, off-the-shelf (OTS) language mo…