paper-with-me

홈 › Papers

Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts

2026-04-20 · Run Xu, Lu Li, Rongzhao Zhang, Jie Xu arxiv

Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over explicit cultural context, making it difficult to jointly maintain image relevance, contextual appropriateness, and humor quality under a specified cultural background. To address this limitation, we introduce a new multimodal generation task, culture-aware humorous captioning, which requires a model to generate a humorous caption conditioned on both an input image and a target cultural context. Captions generated under different cultural contexts are not expected to share the same surface form, but should remain grounded in similar visual situations or humorous rationales.To support this task, we establish a six-dimensional evaluation framework covering image relevance, contextual fit, semantic richness, reasonableness, humor, and creativity. We further propose a staged alignment framework that first initializes the model with high-resource supervision under the Western cultural context, then performs multi-dimensional preference alignment via judge-based GRPO with a Degradation-aware Prototype Repulsion Constraint to mitigate reward hacking in open-ended generation, and finally adapts the model to the Eastern cultural context with a small amount of supervision. Experimental results show that our method achieves stronger overall performance under the proposed evaluation framework, with particularly large gains in contextual fit and a better balance between image relevance and humor under cultural constraints.

📄 PDF Abstract BibTeX arXiv:2604.18091

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generation

Similar Papers 제목 키워드 기반

Hummus: A Dataset of Humorous Multimodal Metaphor Use

2025-04-03 · Xiaoyu Tong, Zhi Zhang, Martha Lewis, Ekaterina Shutova

Metaphor and humor share a lot of common ground, and metaphor is one of the most common humorous mechanisms. This study focuses on the humorous capacity of multimodal metaphors, which has not received due attention in th…

Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning

2024-06-15 · Jifan Zhang, Lalit Jain, Yang Guo, Jiayi Chen 외

We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2 million captions, collected through crowdsourcing rating data for The New Yorker's weekly…

Caption Generation

When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and its Intensity

2022-11-03 · COLING 2022 10 · Khalid Alnajjar, Mika Hämäläinen, Jörg Tiedemann, Jorma Laaksonen 외

Prerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatically detecting humor in the Friends TV sho…

Exploiting Image–Text Synergy for Contextual Image Captioning

2021-04-01 · EACL (LANTERN) 2021 4 · Sreyasi Nag Chowdhury, Rajarshi Bhowmik, Hareesh Ravi, Gerard de Melo 외

Modern web content - news articles, blog posts, educational resources, marketing brochures - is predominantly multimodal. A notable trait is the inclusion of media such as images placed at meaningful locations within a t…

ArticlesImage CaptioningMarketing

HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation

2025-11-21 · Jiajun Zhang, Shijia Luo, Ruikang Zhang, Qi Su arxiv

Humor, as both a creative human activity and a social binding mechanism, has long posed a major challenge for AI generation. Although producing humor requires complex cognitive reasoning and social understanding, theorie…

Image CaptioningSemantic Parsing