paper-with-me

Papers

Role-Play Paradox in Large Language Models: Reasoning Performance Gains and Ethical Dilemmas

2024-09-21 · Jinman Zhao, Zifan Qian, Linbo Cao, Yining Wang, Yitian Ding, Yulan Hu, Zeyu Zhang, Zeyong Jin

Role-play in large language models (LLMs) enhances their ability to generate contextually relevant and high-quality responses by simulating diverse cognitive perspectives. However, our study identifies significant risks associated with this technique. First, we demonstrate that autotuning, a method used to auto-select models' roles based on the question, can lead to the generation of harmful outputs, even when the model is tasked with adopting neutral roles. Second, we investigate how different roles affect the likelihood of generating biased or harmful content. Through testing on benchmarks containing stereotypical and harmful questions, we find that role-play consistently amplifies the risk of biased outputs. Our results underscore the need for careful consideration of both role simulation and tuning processes when deploying LLMs in sensitive or high-stakes contexts.

📄 PDF Abstract BibTeX arXiv:2409.13979

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Reasoning Does Not Necessarily Improve Role-Playing Ability

2025-02-24 · Xiachong Feng, Longxu Dou, Lingpeng Kong

The application of role-playing large language models (LLMs) is rapidly expanding in both academic and commercial domains, driving an increasing demand for high-precision role-playing models. Simultaneously, the rapid ad…

EMNLP: Educator-role Moral and Normative Large Language Models Profiling

2025-08-21 · Yilin Jiang, Mingzi Zhang, Sheng Jin, Zengyi Yu 외 arxiv

Simulating Professions (SP) enables Large Language Models (LLMs) to emulate professional roles. However, comprehensive psychological and ethical evaluation in these contexts remains lacking. This paper introduces EMNLP, …

The Rosetta Paradox: Domain-Specific Performance Inversions in Large Language Models

2024-12-09 · Basab Jha, Ujjwal Puri

While large language models, such as GPT and BERT, have already demonstrated unprecedented skills in everything from natural language processing to domain-specific applications, there came an unexplored phenomenon we ter…

Common Sense ReasoningSpecificity

The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs

2026-05-30 · Zihan Chen, Yiming Zhang, Wenxiang Geng, Zenghui Ding 외 arxiv

Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) frequently exhibit a critical failure mode: they achieve high performance on in-distribution benchmarks while demonstrating brittle reaso…

Reinforcement Learning

Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?

2026-01-11 · Jie Zhu, Yiyang Su, Xiaoming Liu arxiv

Multi-modal large language models (MLLMs) exhibit strong general-purpose capabilities, yet still struggle on Fine-Grained Visual Classification (FGVC), a core perception task that requires subtle visual discrimination an…