paper-with-me

홈 › Papers

Gemma Needs Help: Investigating and Mitigating Emotional Instability in LLMs

2026-02-17 · Anna Soligo, Vladimir Mikulik, William Saunders arxiv

Large language models can generate responses that resemble emotional distress, and this raises concerns around model reliability and safety. We introduce a set of evaluations to investigate expressions of distress in LLMs, and find that these surface emotional instability in Gemma and Gemini models, but not in other families. We find evidence that this difference arises in post-training. Base models from different families (Gemma, Qwen and OLMo) show similar propensities for expressing distress. However, instruct-tuned Gemma expresses substantially more distress than its base model, whereas instruct-tuned Qwen and OLMo express less. We find a simple mitigation for this: direct preference optimisation on just 280 preference pairs reduces Gemma's high-frustration responses from 35% to 0.3% in our evaluations, generalising across question types, user tones, and conversation lengths, without affecting capabilities. These findings show that emotional instability is an issue in some LLMs. We present (1) evaluations to track this behaviour, and (2) a mitigation without downsides in Gemma, with the caveat that upstream training modifications to improve emotional robustness would be significantly better than this post-hoc fix.

📄 PDF Abstract BibTeX arXiv:2603.10011

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mitigating Unhelpfulness in Emotional Support Conversations with Multifaceted AI Feedback

2024-01-11 · Jiashuo Wang, Chunpu Xu, Chak Tou Leong, Wenjie Li 외

An emotional support conversation system aims to alleviate users' emotional distress and assist them in addressing their challenges. To generate supportive responses, it is critical to consider multiple factors such as e…

Contrastive Learning

From Generic Empathy to Personalized Emotional Support: A Self-Evolution Framework for User Preference Alignment

2025-05-22 · Jing Ye, Lu Xiang, Yaping Zhang, Chengqing Zong

Effective emotional support hinges on understanding users' emotions and needs to provide meaningful comfort during multi-turn interactions. Large Language Models (LLMs) show great potential for expressing empathy; howeve…

Consistency of Responses and Continuations Generated by Large Language Models on Social Media

2025-01-14 · Wenlu Fan, Yuqi Zhu, Chenyang Wang, Bin Wang 외

Large Language Models (LLMs) demonstrate remarkable capabilities in text generation, yet their emotional consistency and semantic coherence in social media contexts remain insufficiently understood. This study investigat…

Semantic SimilaritySemantic Textual SimilarityText Generation

Mitigating Exaggerated Safety in Large Language Models

2024-05-08 · Ruchira Ray, Ruchi Bhalani

As the popularity of Large Language Models (LLMs) grow, combining model safety with utility becomes increasingly important. The challenge is making sure that LLMs can recognize and decline dangerous prompts without sacri…

Decision MakingNavigate

Investigating cybersecurity incidents using large language models in latest-generation wireless networks

2025-04-14 · Leonid Legashev, Arthur Zhigalov

The purpose of research: Detection of cybersecurity incidents and analysis of decision support and assessment of the effectiveness of measures to counter information security threats based on modern generative models. Th…

Binary ClassificationData PoisoningFeature ImportanceLarge Language Model