paper-with-me

Papers

Crowd Score: A Method for the Evaluation of Jokes using Large Language Model AI Voters as Judges

2022-12-21 · Fabricio Goes, Zisen Zhou, Piotr Sawicki, Marek Grzes, Daniel G. Brown

This paper presents the Crowd Score, a novel method to assess the funniness of jokes using large language models (LLMs) as AI judges. Our method relies on inducing different personalities into the LLM and aggregating the votes of the AI judges into a single score to rate jokes. We validate the votes using an auditing technique that checks if the explanation for a particular vote is reasonable using the LLM. We tested our methodology on 52 jokes in a crowd of four AI voters with different humour types: affiliative, self-enhancing, aggressive and self-defeating. Our results show that few-shot prompting leads to better results than zero-shot for the voting question. Personality induction showed that aggressive and self-defeating voters are significantly more inclined to find more jokes funny of a set of aggressive/self-defeating jokes than the affiliative and self-enhancing voters. The Crowd Score follows the same trend as human judges by assigning higher scores to jokes that are also considered funnier by human judges. We believe that our methodology could be applied to other creative domains such as story, poetry, slogans, etc. It could both help the adoption of a flexible and accurate standard approach to compare different work in the CC community under a common metric and by minimizing human participation in assessing creative artefacts, it could accelerate the prototyping of creative artefacts and reduce the cost of hiring human participants to rate creative artefacts.

📄 PDF Abstract BibTeX arXiv:2212.11214

Code (1)

creapar/crowdscore 공식 구현

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

ChatGPT is fun, but it is not funny! Humor is still challenging Large Language Models

2023-06-07 · Sophie Jentzsch, Kristian Kersting

Humor is a central aspect of human communication that has not been solved for artificial agents so far. Large language models (LLMs) are increasingly able to capture implicit and contextual information. Especially, OpenA…

valid

Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models

2025-10-21 · Atharvan Dogra, Soumya Suvra Ghosal, Ameet Deshpande, Ashwin Kalyan 외 arxiv

Large language models are increasingly used for creative writing and engagement content, raising safety concerns about the outputs. Therefore, casting humor generation as a testbed, this work evaluates how funniness opti…

Witscript 2: A System for Generating Improvised Jokes Without Wordplay

2023-02-03 · Joe Toplyn

A previous paper presented Witscript, a system for generating conversational jokes that rely on wordplay. This paper extends that work by presenting Witscript 2, which uses a large language model to generate conversation…

ChatbotCommon Sense ReasoningLanguage ModelingLanguage Modelling+1

Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes

2025-07-17 · Tyler Loakman, William Thorne, Chenghua Lin

Humour, as a complex language form, is derived from myriad aspects of life, whilst existing work on computational humour has focussed almost exclusively on short pun-based jokes. In this work, we investigate whether the …

Common Sense ReasoningWorld Knowledge

Witscript: A System for Generating Improvised Jokes in a Conversation

2023-02-03 · Joe Toplyn

A chatbot is perceived as more humanlike and likeable if it includes some jokes in its output. But most existing joke generators were not designed to be integrated into chatbots. This paper presents Witscript, a novel jo…

ChatbotLanguage ModelingLanguage ModellingSentence