paper-with-me

홈 › Papers

SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses

2026-07-02 · Anna Chorna arxiv

Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abilities in these circumstances remain underexplored. Existing benchmarks emphasize multilingual performance but rarely examine crisis-related empathy and cultural grounding in low-to-mid-resource languages. We introduce SPLIT, a 500-prompt benchmark designed to evaluate LLM consistency in generating emotionally grounded responses across five categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. We evaluate three technically diverse LLMs across three dimensions: Empathetic Accuracy, Linguistic Naturalness, and Contextual & Cultural Grounding. The framework aims to assess and compare the quality of LLM responses in both English and Ukrainian languages, as well as to explore the reliability of the LLM-as-a-jury paradigm. Our findings reveal that Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct degrade when transitioning to Ukrainian, while DeepSeek-V3 remains comparatively stable within our benchmark. We additionally find that human and AI evaluators agree weakly on empathy and naturalness but diverge on cultural grounding. We further argue that producing Ukrainian text is not equivalent to producing Ukrainian emotional support. Our findings may assist in the future development of more culturally tailored benchmark designs, as well as encourage a stronger emphasis on human-centered evaluation.

📄 PDF Abstract BibTeX arXiv:2607.02049

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Same Lesson, Different Story: Cross-Lingual Reconstruction of Cultural Narratives in Large Language Models

2026-06-23 · Jory Alshaalan, Haya Albaker, Abeer Aldayel, Aljawharah Alabdullatif 외 arxiv

The evaluation of cultural grounding context becomes complex when multiple cultures convey the same moral lesson. This challenge is particularly relevant to large language models (LLMs), which produce narratives across a…

Semantic Similarity

Cultural Prompting Improves the Empathy and Cultural Responsiveness of GPT-Generated Therapy Responses

2025-10-19 · Serena Jinchen Xie, Shumenghui Zhai, Yanjing Liang, Jingyi Li 외 arxiv

Large Language Model (LLM)-based conversational agents offer promising solutions for mental health support, but lack cultural responsiveness for diverse populations. This study evaluated the effectiveness of cultural pro…

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

2026-08-24 · Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha arxiv

Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording …

MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation

2025-10-07 · Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu 외 arxiv

Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a comprehensive framework to evaluate LLMs…

Visual Question Answering

SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing

2026-08-22 · Yifan Lyu, Dianqing Lin, Xinran Li, Jiaqi Qiao 외 arxiv

Hate speech research has moved from coarse-grained classification towards structured parsing, where systems jointly identify targets, supporting arguments, and target-level labels. Documents with multiple targets, confli…