paper-with-me

Papers

Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems

2025-09-28 · Guojian Li, Chengyou Wang, Hongfei Xue, Shuiyuan Wang, Dehui Gao, Zihan Zhang, Yuke Lin, Wenjie Li, Longshuai Xiao, Zhonghua Fu, Lei Xie arxiv

Full-duplex interaction is crucial for natural human-machine communication, yet remains challenging as it requires robust turn-taking detection to decide when the system should speak, listen, or remain silent. Existing solutions either rely on dedicated turn-taking models, most of which are not open-sourced. The few available ones are limited by their large parameter size or by supporting only a single modality, such as acoustic or linguistic. Alternatively, some approaches finetune LLM backbones to enable full-duplex capability, but this requires large amounts of full-duplex data, which remain scarce in open-source form. To address these issues, we propose Easy Turn, an open-source, modular turn-taking detection model that integrates acoustic and linguistic bimodal information to predict four dialogue turn states: complete, incomplete, backchannel, and wait, accompanied by the release of Easy Turn trainset, a 1,145-hour speech dataset designed for training turn-taking detection models. Compared to existing open-source models like TEN Turn Detection and Smart Turn V2, our model achieves state-of-the-art turn-taking detection accuracy on our open-source Easy Turn testset. The data and model will be made publicly available on GitHub.

📄 PDF Abstract BibTeX arXiv:2509.23938

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNs

2018-08-31 · Matthew Roddy, Gabriel Skantze, Naomi Harte

In human conversational interactions, turn-taking exchanges can be coordinated using cues from multiple modalities. To design spoken dialog systems that can conduct fluid interactions it is desirable to incorporate cues …

Prediction

Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals

2025-05-19 · Yuxin Lin, Yinglin Zheng, Ming Zeng, Wangzheng Shi

This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets,…

Modeling Turn-Taking with Semantically Informed Gestures

2025-10-22 · Varsha Suresh, M. Hamza Mughal, Christian Theobalt, Vera Demberg arxiv

In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide complementary cues for modeling these trans…

Acoustic and Machine Learning Methods for Speech-Based Suicide Risk Assessment: A Systematic Review

2025-05-20 · Ambre Marie, Marine Garnier, Thomas Bertin, Laura Machart 외

Suicide remains a public health challenge, necessitating improved detection methods to facilitate timely intervention and treatment. This systematic review evaluates the role of Artificial Intelligence (AI) and Machine L…

Articles

Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR

2024-09-03 · Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous fea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)cross-modal alignmentspeech-recognition+2