paper-with-me

홈 › Papers

Oogiri-Master: Benchmarking Humor Understanding via Oogiri

2025-12-25 · Soichiro Murakami, Hidetaka Kamigaito, Hiroya Takamura, Manabu Okumura arxiv

Humor is a salient testbed for human-like creative thinking in large language models (LLMs). We study humor using the Japanese creative response game Oogiri, in which participants produce witty responses to a given prompt, and ask the following research question: What makes such responses funny to humans? Previous work has offered only limited reliable means to answer this question. Existing datasets contain few candidate responses per prompt, expose popularity signals during ratings, and lack objective and comparable metrics for funniness. Thus, we introduce Oogiri-Master and Oogiri-Corpus, which are a benchmark and dataset designed to enable rigorous evaluation of humor understanding in LLMs. Each prompt is paired with approximately 100 diverse candidate responses, and funniness is rated independently by approximately 100 human judges without access to others' ratings, reducing popularity bias and enabling robust aggregation. Using Oogiri-Corpus, we conduct a quantitative analysis of the linguistic factors associated with funniness, such as text length, ambiguity, and incongruity resolution, and derive objective metrics for predicting human judgments. Subsequently, we benchmark a range of LLMs and human baselines in Oogiri-Master, demonstrating that state-of-the-art models approach human performance and that insight-augmented prompting improves the model performance. Our results provide a principled basis for evaluating and advancing humor understanding in LLMs.

📄 PDF Abstract BibTeX arXiv:2512.21494

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation

2023-12-05 · CVPR 2024 1 · Shanshan Zhong, Zhongzhan Huang, ShangHua Gao, Wushao Wen 외

Chain-of-Thought (CoT) guides large language models (LLMs) to reason step-by-step, and can motivate their logical reasoning ability. While effective for logical tasks, CoT is not conducive to creative problem-solving whi…

Logical Reasoning

Assessing the Capabilities of LLMs in Humor:A Multi-dimensional Analysis of Oogiri Generation and Evaluation

2025-11-12 · Ritsu Sakabe, Hwichan Kim, Tosho Hirasawa, Mamoru Komachi arxiv

Computational humor is a frontier for creating advanced and engaging natural language processing (NLP) applications, such as sophisticated dialogue systems. While previous studies have benchmarked the humor capabilities …

Who Laughs with Whom? Disentangling Influential Factors in Humor Preferences across User Clusters and LLMs

2026-01-06 · Soichiro Murakami, Hidetaka Kamigaito, Hiroya Takamura, Manabu Okumura arxiv

Humor preferences vary widely across individuals and cultures, complicating the evaluation of humor using large language models (LLMs). In this study, we model heterogeneity in humor preferences in Oogiri, a Japanese cre…

HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation

2025-11-21 · Jiajun Zhang, Shijia Luo, Ruikang Zhang, Qi Su arxiv

Humor, as both a creative human activity and a social binding mechanism, has long posed a major challenge for AI generation. Although producing humor requires complex cognitive reasoning and social understanding, theorie…

Image CaptioningSemantic Parsing

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models

2025-01-25 · Zhongzhan Huang, Shanshan Zhong, Pan Zhou, ShangHua Gao 외

Recently, numerous benchmarks have been developed to evaluate the logical reasoning abilities of large language models (LLMs). However, assessing the equally important creative capabilities of LLMs is challenging due to …

Logical Reasoning