paper-with-me

홈 › Papers

Leveraging Real Conversational Data for Multi-Channel Continuous Speech Separation

2022-04-07 · Xiaofei Wang, Dongmei Wang, Naoyuki Kanda, Sefik Emre Eskimez, Takuya Yoshioka

Existing multi-channel continuous speech separation (CSS) models are heavily dependent on supervised data - either simulated data which causes data mismatch between the training and real-data testing, or the real transcribed overlapping data, which is difficult to be acquired, hindering further improvements in the conversational/meeting transcription tasks. In this paper, we propose a three-stage training scheme for the CSS model that can leverage both supervised data and extra large-scale unsupervised real-world conversational data. The scheme consists of two conventional training approaches -- pre-training using simulated data and ASR-loss-based training using transcribed data -- and a novel continuous semi-supervised training between the two, in which the CSS model is further trained by using real data based on the teacher-student learning framework. We apply this scheme to an array-geometry-agnostic CSS model, which can use the multi-channel data collected from any microphone array. Large-scale meeting transcription experiments are carried out on both Microsoft internal meeting data and the AMI meeting corpus. The steady improvement by each training stage has been observed, showing the effect of the proposed method that enables leveraging real conversational data for CSS model training.

📄 PDF Abstract BibTeX arXiv:2204.03232

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

2023-11-01 · Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi 외

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech-to-Text+2

Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection

2024-10-21 · Koji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara

In human conversations, short backchannel utterances such as "yeah" and "oh" play a crucial role in facilitating smooth and engaging dialogue. These backchannels signal attentiveness and understanding without interruptin…

PredictionType prediction

M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses

2024-09-17 · Yufeng Yang, Desh Raj, Ju Lin, Niko Moritz 외

The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tas…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+3

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations

2026-04-25 · Bhaskar Singh, Shobhit Banga, Mahima Manik, Pranav Sharma arxiv

Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first ope…

Text Generation

Multi-channel Conversational Speaker Separation via Neural Diarization

2023-11-15 · Hassan Taherian, DeLiang Wang

When dealing with overlapped speech, the performance of automatic speech recognition (ASR) systems substantially degrades as they are designed for single-talker speech. To enhance ASR performance in conversational or mee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Separationspeech-recognition+1