paper-with-me

홈 › Papers

Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations

2025-09-27 · Yihao Wu, Tianrui Wang, Yizhou Peng, Yi-Wen Chao, Xuyi Zhuang, Xinsheng Wang, Shunshun Yin, Ziyang Ma arxiv

While biases in large language models (LLMs), such as stereotypes and cultural tendencies in outputs, have been examined and identified, their presence and characteristics in spoken dialogue models (SDMs) with audio input and output remain largely unexplored. Paralinguistic features, such as age, gender, and accent, can affect model outputs; when compounded by multi-turn conversations, these effects may exacerbate biases, with potential implications for fairness in decision-making and recommendation tasks. In this paper, we systematically evaluate biases in speech LLMs and study the impact of multi-turn dialogues with repeated negative feedback. Bias is measured using Group Unfairness Score (GUS) for decisions and similarity-based normalized statistics rate (SNSR) for recommendations, across both open-source models like Qwen2.5-Omni and GLM-4-Voice, as well as closed-source APIs such as GPT-4o Audio and Gemini-2.5-Flash. Our analysis reveals that closed-source models generally exhibit lower bias, while open-source models are more sensitive to age and gender, and recommendation tasks tend to amplify cross-group disparities. We found that biased decisions may persist in multi-turn conversations. This work provides the first systematic study of biases in end-to-end spoken dialogue models, offering insights towards fair and reliable audio-based interactive systems. To facilitate further research, we release the FairDialogue dataset and evaluation code.

📄 PDF Abstract BibTeX arXiv:2510.02352

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WavChat: A Survey of Spoken Dialogue Models

2024-11-15 · Shengpeng Ji, Yifu Chen, Minghui Fang, Jialong Zuo 외

Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier cascaded spoken dialogue models that compris…

speech-recognitionSpeech RecognitionSpoken Dialogue SystemsSurvey+2

The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era

2026-01-09 · Zhixian Zhao, Shuiyuan Wang, Guojian Li, Hongfei Xue 외 arxiv

Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and huma…

Emotional Intelligence

Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents

2024-09-23 · Bandhav Veluri, Benjamin N Peloquin, Bokai Yu, Hongyu Gong 외

Despite broad interest in modeling spoken dialogue agents, most approaches are inherently "half-duplex" -- restricted to turn-based interaction with responses requiring explicit prompting by the user or implicit tracking…

2k

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

2026-05-26 · Heriberto Cuayahuitl, Grace Jang arxiv

Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, their application to textual or spoken medical consultations is still a…

Data Augmentation

Are LLMs Robust for Spoken Dialogues?

2024-01-04 · Seyed Mahed Mousavi, Gabriel Roccabruna, Simone Alghisi, Massimo Rizzoli 외

Large Pre-Trained Language Models have demonstrated state-of-the-art performance in different downstream tasks, including dialogue state tracking and end-to-end response generation. Nevertheless, most of the publicly ava…

Dialogue State TrackingResponse Generation