paper-with-me

홈 › Papers

Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning

2026-06-26 · Austin MY Cheung, Yi Yang arxiv

Recent work has shown that fine-tuning large language models (LLMs) for social warmth degrades factual reliability and increases sycophancy. We investigate a related but distinct failure mode: warmth fine-tuning also weakens adversarial safety, making models more susceptible to jailbreaks and harmful output generation. We examine whether this reflects an inherent consequence of empathetic adaptation or an artifact of data construction. To address this, we introduce a persona-driven rewriting pipeline that conditions user turns on low agreeableness and pairs this with warm, de-escalating assistant responses. Across three experiments on four models, our approach reduces jailbreak susceptibility and harmful output rates relative to generic warmth fine-tuning baselines, while preserving conversational warmth. Representational probing provides suggestive evidence that this conditioning reduces the geometric alignment between warmth and compliance directions in latent space. These results show that safer empathetic fine-tuning is achievable through data design alone, without safety labels, harm detectors, or changes to the training objective.

📄 PDF Abstract BibTeX arXiv:2606.27709

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bullying the Machine: How Personas Increase LLM Vulnerability

2025-05-19 · Ziwei Xu, Udit Sanghi, Mohan Kankanhalli

Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects model safety under bullying, an adversar…

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models

2026-04-12 · Arya Shah, Deepali Mishra, Chaklam Silpasuwanchai arxiv

Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while valuable, raises concerns about sycophancy: the tendency to provide r…

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

2026-07-29 · Hasibur Rahman, Smit Desai arxiv

Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable a…

What makes your model a low-empathy or warmth person: Exploring the Origins of Personality in LLMs

2024-10-07 · Shu Yang, Shenzhe Zhu, Ruoxuan Bao, Liang Liu 외

Large language models (LLMs) have demonstrated remarkable capabilities in generating human-like text and exhibiting personality traits similar to those in humans. However, the mechanisms by which LLMs encode and express …

Persona Cartography: Charting Language Model Personality Traits in Weight Space

2026-07-08 · Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk, Irakli Shalibashvili 외 arxiv

Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them. Our central insight is to tre…