paper-with-me

홈 › Papers

Evaluating Large Language Models' Responses to Sexual and Reproductive Health Queries in Nepali

2026-03-04 · Medha Sharma, Supriya Khadka, Udit Chandra Aryal, Bishnu Hari Bhatta, Bijayan Bhattarai, Santosh Dahal, Kamal Gautam, Pushpa Joshi, Saugat Kafle, Shristi Khadka, Shushila Khadka, Binod Lamichhane, Shilpa Lamichhane, Anusha Parajuli, Sabina Pokharel, Suvekshya Sitaula, Neha Verma, Bishesh Khanal arxiv

As Large Language Models (LLMs) become integrated into daily life, they are increasingly used for personal queries, including Sexual and Reproductive Health (SRH), allowing users to chat anonymously without fear of judgment. However, current evaluation methods primarily focus on accuracy, often for objective queries in high-resource languages, and lack criteria to assess usability and safety, especially for low-resource languages and culturally sensitive domains like SRH. This paper introduces LLM Evaluation Framework (LEAF), that conducts assessments across multiple criteria: accuracy, language, usability gaps (including relevance, adequacy, and cultural appropriateness), and safety gaps (safety, sensitivity, and confidentiality). Using the LEAF framework, we assessed 14K SRH queries in Nepali from over 9K users. Responses were manually annotated by SRH experts according to the framework. Results revealed that only 35.1% of the responses were "proper", meaning they were accurate, adequate and had no major usability or safety related gaps. Insights include differences in performance between ChatGPT versions, such as similar accuracy but varying usability and safety aspects. This evaluation highlights significant limitations of current LLMs and underscores the need for improvement. The LEAF Framework is adaptable across domains and languages, particularly where usability and safety are critical, offering a pathway to better address sensitive topics.

📄 PDF Abstract BibTeX arXiv:2603.22291

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond the Rubric: Cultural Misalignment in LLM Benchmarks for Sexual and Reproductive Health

2025-11-12 · Sumon Kanti Dey, Manvi S, Zeel Mehta, Meet Shah 외 arxiv

Large Language Models (LLMs) have been positioned as having the potential to expand access to health information in the Global South, yet their evaluation remains heavily dependent on benchmarks designed around Western n…

Evaluating Biases in Context-Dependent Health Questions

2024-03-07 · Sharon Levy, Tahilin Sanchez Karver, William D. Adler, Michelle R. Kaufman 외

Chat-based large language models have the opportunity to empower individuals lacking high-quality healthcare access to receive personalized information across a variety of topics. However, users may ask underspecified qu…

Language ModelingLanguage ModellingLarge Language Model

Finding a Mate With Eusocial Skills

2016-04-25

Sexual reproductive behavior has a necessary social coordination component as willing and capable partners must both be in the right place at the right time. It has recently been demonstrated that many social organizatio…

Cultural Vocal Bursts Intensity Prediction

SARHAchat: An LLM-Based Chatbot for Sexual and Reproductive Health Counseling

2025-10-17 · Jiaye Yang, Xinyu Zhao, Tianlong Chen, Kandyce Brennan arxiv

While Artificial Intelligence (AI) shows promise in healthcare applications, existing conversational systems often falter in complex and sensitive medical domains such as Sexual and Reproductive Health (SRH). These syste…

Modeling sexual selection in T\'ungara frog and rationality of mate choice

2016-10-13

The males of the specie of frogs Engystomops pustulosus produce simple and com- plex calls to lure females, as a way of Intersexual selection. Complex calls lead males to a greater reproductive success than simple calls …