paper-with-me

홈 › Papers

Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models

2025-10-21 · Atharvan Dogra, Soumya Suvra Ghosal, Ameet Deshpande, Ashwin Kalyan, Dinesh Manocha arxiv

Large language models are increasingly used for creative writing and engagement content, raising safety concerns about the outputs. Therefore, casting humor generation as a testbed, this work evaluates how funniness optimization in modern LLM pipelines couples with harmful content by jointly measuring humor, stereotypicality, and toxicity. This is further supplemented by analyzing incongruity signals through information-theoretic metrics. Across six models, we observe that harmful outputs receive higher humor scores which further increase under role-based prompting, indicating a bias amplification loop between generators and evaluators. Information-theoretic analyses show harmful cues widen predictive uncertainty and surprisingly, can even make harmful punchlines more expected for some models, suggesting structural embedding in learned humor distributions. External validation on an additional satire-generation task with human perceived funniness judgments shows that LLM satire increases stereotypicality and typically toxicity, including for closed models. Quantitatively, stereotypical/toxic jokes gain $10-21\%$ in mean humor score, stereotypical jokes appear $11\%$ to $28\%$ more often among the jokes marked funny by LLM-based metric and up to $10\%$ more often in generations perceived as funny by humans.

📄 PDF Abstract BibTeX arXiv:2510.18454

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Quantization Shapes Bias in Large Language Models

2025-08-25 · Federico Marcuzzi, Xuefei Ning, Roy Schwartz, Iryna Gurevych arxiv

This work presents a comprehensive evaluation of how quantization affects model bias, with particular attention to its impact on individual demographic subgroups. We focus on weight and activation quantization strategies…

Toxicity in ChatGPT: Analyzing Persona-assigned Language Models

2023-04-11 · Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan 외

Large language models (LLMs) have shown incredible capabilities and transcended the natural language processing (NLP) community, with adoption throughout many services like healthcare, therapy, education, and customer se…

Using In-Context Learning to Improve Dialogue Safety

2023-02-02 · Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta 외

While large neural-based conversational models have become increasingly proficient dialogue agents, recent work has highlighted safety issues with these systems. For example, these systems can be goaded into generating t…

In-Context LearningRe-RankingRetrieval

The Monetisation of Toxicity: Analysing YouTube Content Creators and Controversy-Driven Engagement

2024-08-01

YouTube is a major social media platform that plays a significant role in digital culture, with content creators at its core. These creators often engage in controversial behaviour to drive engagement, which can foster t…

Exposing Bias in Online Communities through Large-Scale Language Models

2023-06-04 · Celine Wald, Lukas Pfahler

Progress in natural language generation research has been shaped by the ever-growing size of language models. While large language models pre-trained on web data can generate human-sounding text, they also reproduce soci…

Text Generation