paper-with-me

홈 › Papers

Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions

2026-06-25 · Abla Bedoui, Ashley L. Greene, Mohammed Cherkaoui arxiv

Large language models (LLMs) are increasingly being integrated into mental health support tools and other psychologically sensitive conversational applications. In such settings, behavioral stability and consistency are important for trustworthy human-AI interaction. However, semantically similar concerns can be presented through different contextual framings, potentially eliciting different model responses. Such framing-sensitive variability may challenge user expectations regarding system behavior and complicate the assessment of AI reliability. While prior studies have primarily examined such effects at the behavioral level, less is known about how framing-related variation is reflected in the internal representations of aligned language models. In this work, we investigate these effects using controlled matched prompts spanning multiple contextual framing conditions across several instruction-tuned model families. Across architectures, framing systematically alters interpretive response tendencies. Layer-wise probing analyses show that behavior-associated information remains decodable throughout transformer depth, with architecture-dependent variation in decoding strength. Moreover, held-out framing probes remained consistently above chance across architectures despite strong lexical baselines. Activation steering experiments further suggest that framing-associated representational directions can partially modulate downstream behavioral outcomes. Finally, these findings indicate that robustness to contextual variation may represent an important consideration when evaluating the consistency and trustworthiness of conversational AI systems deployed in mental-health-oriented interactions.

📄 PDF Abstract BibTeX arXiv:2606.26982

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models

2025-12-15 · TK Lee arxiv

Large language models (LLMs) are widely deployed as general-purpose tools, yet extended interaction can reveal behavioral patterns not captured by standard quantitative benchmarks. We present a qualitative case-study met…

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

2026-01-29 · Katherine Elkins, Jon Chun arxiv

Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framing. We audit 16 models across 14 ethically fraught dilemmas using polarit…

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

2026-08-25 · Augusto Camargo arxiv

Evaluations of generative language models frequently interpret observable behavioral traits, such as political stance, brand inclination, and normative framing, as manifestations of model weights, post-training alignment…

Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores

2026-04-23 · Shevya Panda, Shinjini Bose, Ananya Joshi arxiv

Large language models (LLMs) are increasingly utilized in clinical reasoning and risk assessment. However, their interpretive reliability in critical and indeterminate domains such as psychiatry remains unclear. Prior wo…

Interactional Fairness in LLM Multi-Agent Systems: An Evaluation Framework

2025-05-17 · Ruta Binkyte

As large language models (LLMs) are increasingly used in multi-agent systems, questions of fairness should extend beyond resource distribution and procedural design to include the fairness of how agents communicate. Draw…

Fairness