paper-with-me

홈 › Papers

Large Language Models Know What To Say But Not When To Speak

2024-10-21 · Muhammad Umair, Vasanth Sarathy, Jp de Ruiter

Turn-taking is a fundamental mechanism in human communication that ensures smooth and coherent verbal interactions. Recent advances in Large Language Models (LLMs) have motivated their use in improving the turn-taking capabilities of Spoken Dialogue Systems (SDS), such as their ability to respond at appropriate times. However, existing models often struggle to predict opportunities for speaking -- called Transition Relevance Places (TRPs) -- in natural, unscripted conversations, focusing only on turn-final TRPs and not within-turn TRPs. To address these limitations, we introduce a novel dataset of participant-labeled within-turn TRPs and use it to evaluate the performance of state-of-the-art LLMs in predicting opportunities for speaking. Our experiments reveal the current limitations of LLMs in modeling unscripted spoken interactions, highlighting areas for improvement and paving the way for more naturalistic dialogue systems.

📄 PDF Abstract BibTeX arXiv:2410.16044

Code (0)

등록된 구현이 없습니다.

Tasks

Spoken Dialogue Systems

Similar Papers 제목 키워드 기반

M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models

2025-10-22 · Yejin Kwon, Taewoo Kang, Hyunsoo Yoon, Changouk Kim arxiv

We present M3-SLU, a new multimodal large language model (MLLM) benchmark for evaluating multi-speaker, multi-turn spoken language understanding. While recent models show strong performance in speech and text comprehensi…

Spoken Language UnderstandingQuestion Answering

HumanOmni-Speaker: Identifying Who said What and When

2026-03-23 · Detao Bai, Zhiheng Ma, Xihan Wei arxiv

While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, multi-person conversational dynamics to accu…

Natural Language QueriesSpeaker Diarization

Beyond Words: Multimodal LLM Knows When to Speak

2025-05-20 · Zikai Liao, Yi Ouyang, Yi-Lun Lee, Chen-Ping Yu 외

While large language model (LLM)-based chatbots have demonstrated strong capabilities in generating coherent and contextually relevant responses, they often struggle with understanding when to speak, particularly in deli…

Large Language Model

What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification

2026-03-25 · Massa Baali, Sarthak Bisht, Rita Singh, Bhiksha Raj arxiv

Speaker verification at large scale remains an open challenge as fixed-margin losses treat all samples equally regardless of quality. We hypothesize that mislabeled or degraded samples introduce noisy gradients that disr…

Speaker Verification

What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects

2024-02-19 · Verena Blaschke, Christoph Purschke, Hinrich Schütze, Barbara Plank

Natural language processing (NLP) has largely focused on modelling standardized languages. More recently, attention has increasingly shifted to local, non-standardized languages and dialects. However, the relevant speake…

Machine Translation