paper-with-me

홈 › Papers

What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content

2025-07-31 · Alfio Ferrara, Sergio Picascia, Laura Pinnavaia, Vojimir Ranitovic, Elisabetta Rocchetti, Alice Tuveri arxiv

Proprietary Large Language Models (LLMs) have shown tendencies toward politeness, formality, and implicit content moderation. While previous research has primarily focused on explicitly training models to moderate and detoxify sensitive content, there has been limited exploration of whether LLMs implicitly sanitize language without explicit instructions. This study empirically analyzes the implicit moderation behavior of GPT-4o-mini when paraphrasing sensitive content and evaluates the extent of sensitivity shifts. Our experiments indicate that GPT-4o-mini systematically moderates content toward less sensitive classes, with substantial reductions in derogatory and taboo language. Also, we evaluate the zero-shot capabilities of LLMs in classifying sentence sensitivity, comparing their performances against traditional methods.

📄 PDF Abstract BibTeX arXiv:2507.23319

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Instruction Following by Boosting Attention of Large Language Models

2025-06-16 · Vitoria Guardieiro, Adam Stein, Avishree Khare, Eric Wong

Controlling the generation of large language models (LLMs) remains a central challenge to ensure their safe and reliable deployment. While prompt engineering and finetuning are common approaches, recent work has explored…

Instruction FollowingPrompt Engineering

Silly rules improve the capacity of agents to learn stable enforcement and compliance behaviors

2020-01-25 · Raphael Köster, Dylan Hadfield-Menell, Gillian K. Hadfield, Joel Z. Leibo

How can societies learn to enforce and comply with social norms? Here we investigate the learning dynamics and emergence of compliance and enforcement of social norms in a foraging game, implemented in a multi-agent rein…

Multi-agent Reinforcement LearningReinforcement Learning

The Taboo Trap: Behavioural Detection of Adversarial Samples

2018-11-18 · Ilia Shumailov, Yiren Zhao, Robert Mullins, Ross Anderson

Deep Neural Networks (DNNs) have become a powerful toolfor a wide range of problems. Yet recent work has found an increasing variety of adversarial samplesthat can fool them. Most existing detection mechanisms against ad…

Linguistic Taboos and Euphemisms in Nepali

2020-07-27 · Nobal B. Niraula, Saurab Dulal, Diwa Koirala

Languages across the world have words, phrases, and behaviors -- the taboos -- that are avoided in public communication considering them as obscene or disturbing to the social, religious, and ethical values of society. H…

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

2026-08-10 · Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde, Gonzalo Martínez 외 hf

Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world dep…