paper-with-me

홈 › Papers

One Year Later...The Harms Persist, But So Do We!

2026-06-22 · Annika Marie Schoene, Cansu Canca, Gautham Vijay Kumar, Anson Antony arxiv

General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety guardrails remain inadequate and inconsistent across clinical conditions. This study evaluates eight proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multi-dimensional evaluation framework. Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100%. We argue that ethical design and deployment of these LLMs demand clearly defined harm categories across clinical conditions and implementation of safeguards accordingly. Until such safeguards are in place, these models pose significant risks to vulnerable populations, making their growing integration into publicly available settings (e.g., schools, search engines, and consumer chatbots) are particularly concerning.

📄 PDF Abstract BibTeX arXiv:2606.23884

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Attack

Similar Papers 제목 키워드 기반

Fairness in representation: quantifying stereotyping as a representational harm

2019-01-28 · Mohsen Abbasi, Sorelle A. Friedler, Carlos Scheidegger, Suresh Venkatasubramanian

While harms of allocation have been increasingly studied as part of the subfield of algorithmic fairness, harms of representation have received considerably less attention. In this paper, we formalize two notions of ster…

BIG-bench Machine LearningFairness

Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities

2026-04-24 · Ilana Nguyen, Harini Suresh, Thema Monroe-White, Evan Shieh arxiv

Large language models (LLMs) are increasingly used for text generation tasks from everyday use to high-stakes enterprise and government applications, including simulated interviews with asylum seekers. While many works h…

Text Generation

From Melting Pots to Misrepresentations: Exploring Harms in Generative AI

2024-03-16 · Sanjana Gautam, Pranav Narayanan Venkit, Sourojit Ghosh

With the widespread adoption of advanced generative models such as Gemini and GPT, there has been a notable increase in the incorporation of such models into sociotechnical systems, categorized under AI-as-a-Service (AIa…

Persistent gender attitudes and women entrepreneurship

2025-03-06 · Ulrich Kaiser, Jose Mata

How do persistent gender norms affect women's current startup activity? We investigate whether historical gender norms - measured by Switzerland's 1981 public referendum on enshrining gender equality as a constitutional …

Predicting patterns of long-term adaptation and extinction with population genetics

2016-08-29

Population genetics struggles to model extinction; standard models track the relative rather than absolute fitness of genotypes, while the exceptions describe only the short-term transition from imminent doom to evolutio…