paper-with-me

Papers

Mitigating Harmful Erraticism in LLMs Through Dialectical Behavior Therapy Based De-Escalation Strategies

2025-09-06 · Pooja Rangarajan, Jacob Boyle arxiv

The escalating demand for personalized AI chatbot interactions, capable of dynamically adapting to user emotional states and real-time requests, has highlighted critical limitations in current development paradigms. Existing methodologies, which rely on baseline programming, custom personalities, and manual response adjustments, often prove difficult to maintain and are susceptible to errors such as hallucinations, erratic outputs, and software bugs. This paper hypothesizes that a framework rooted in human psychological principles, specifically therapeutic modalities, can provide a more robust and sustainable solution than purely technical interventions. Drawing an analogy to the simulated neural networks of AI mirroring the human brain, we propose the application of Dialectical Behavior Therapy (DBT) principles to regulate chatbot responses to diverse user inputs. This research investigates the impact of a DBT-based framework on AI chatbot performance, aiming to ascertain its efficacy in yielding more reliable, safe, and accurate responses, while mitigating the occurrence of hallucinations, erratic behaviors, and other systemic issues.

📄 PDF Abstract BibTeX arXiv:2510.15889

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models

2024-01-24 · Hongzhan Lin, Ziyang Luo, Wei Gao, Jing Ma 외

The age of social media is flooded with Internet memes, necessitating a clear grasp and effective identification of harmful ones. This task presents a significant challenge due to the implicit meaning embedded in memes, …

Hateful Meme ClassificationLanguage ModellingSmall Language ModelText Generation

Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

2023-10-10 · Zeming Wei, Yifei Wang, Ang Li, Yichuan Mo 외

Large Language Models (LLMs) have shown remarkable success in various tasks, yet their safety and the risk of generating harmful content remain pressing concerns. In this paper, we delve into the potential of In-Context …

In-Context LearningLanguage ModellingSafety Alignment

Learn and Unlearn in Multilingual LLMs

2024-06-19 · Taiming Lu, Philipp Koehn

This paper investigates the propagation of harmful information in multilingual large language models (LLMs) and evaluates the efficacy of various unlearning methods. We demonstrate that fake information, regardless of th…

SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering

2025-08-15 · Utsav Maskey, Sumit Yadav, Mark Dras, Usman Naseem arxiv

LLMs increasingly exhibit over-refusal behavior, where safety mechanisms cause models to reject benign instructions that seemingly resemble harmful content. This phenomenon diminishes utility in production applications t…

Sentiment Analysis

Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms

2025-01-23 · Rajvardhan Oak, Muhammad Haroon, Claire Jo, Magdalena Wojcieszak 외

Social media platforms utilize Machine Learning (ML) and Artificial Intelligence (AI) powered recommendation algorithms to maximize user engagement, which can result in inadvertent exposure to harmful content. Current mo…

Re-Ranking