paper-with-me

홈 › Papers

Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue

2024-09-07 · Junkai Wu, Xulin Fan, Bo-Ru Lu, Xilin Jiang, Nima Mesgarani, Mark Hasegawa-Johnson, Mari Ostendorf

In recent years, we have observed a rapid advancement in speech language models (SpeechLLMs), catching up with humans' listening and reasoning abilities. SpeechLLMs have demonstrated impressive spoken dialog question-answering (SQA) performance in benchmarks like Gaokao, the English listening test of the college entrance exam in China, which seemingly requires understanding both the spoken content and voice characteristics of speakers in a conversation. However, after carefully examining Gaokao's questions, we find the correct answers to many questions can be inferred from the conversation transcript alone, i.e.\ without speaker segmentation and identification. Our evaluation of state-of-the-art models Qwen-Audio and WavLLM on both Gaokao and our proposed "What Do You Like?" dataset shows a significantly higher accuracy in these context-based questions than in identity-critical questions, which can only be answered reliably with correct speaker identification. The results and analysis suggest that when solving SQA, the current SpeechLLMs exhibit limited speaker awareness from the audio and behave similarly to an LLM reasoning from the conversation transcription without sound. We propose that tasks focused on identity-critical questions could offer a more accurate evaluation framework of SpeechLLMs in SQA.

📄 PDF Abstract BibTeX arXiv:2409.04927

Code (1)

wjk0925/slt2024-speechllm-speaker-understanding 공식 구현

Tasks

Question AnsweringSpeaker Identification

Similar Papers 제목 키워드 기반

Leveraging supplementary text data to kick-start automatic speech recognition system development with limited transcriptions

2023-02-09 · Nay San, Martijn Bartelds, Blaine Billings, Ella de Falco 외

Recent research using pre-trained transformer models suggests that just 10 minutes of transcribed speech may be enough to fine-tune such a model for automatic speech recognition (ASR) -- at least if we can also leverage …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Social Robot with Inner Speech for Dietary Guidance

2025-05-13 · Valerio Belcamino, Alessandro Carfì, Valeria Seidita, Fulvio Mastrogiovanni 외

We explore the use of inner speech as a mechanism to enhance transparency and trust in social robots for dietary advice. In humans, inner speech structures thought processes and decision-making; in robotics, it improves …

Computational EfficiencyDecision MakingNatural Language Understanding

Determining Code Words in Euphemistic Hate Speech Using Word Embedding Networks

2018-10-01 · WS 2018 10 · Rijul Magu, Jiebo Luo

While analysis of online explicit abusive language detection has lately seen an ever-increasing focus, implicit abuse detection remains a largely unexplored space. We carry out a study on a subcategory of implicit hate: …

Abuse DetectionAbusive LanguageCommunity DetectionHate Speech Detection+1

Improving Speech Recognition for Indic Languages using Language Model

2022-03-30 · Ankur Dhuriya, Harveen Singh Chadha, Anirudh Gupta, Priyanshi Shah 외

We study the effect of applying a language model (LM) on the output of Automatic Speech Recognition (ASR) systems for Indic languages. We fine-tune wav2vec $2.0$ models for $18$ Indic languages and adjust the results wit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

2025-05-22 · Tianduo Wang, Lu Xu, Wei Lu, Shanbo Cheng

Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+3