paper-with-me

홈 › Papers

Joint speaker diarisation and tracking in switching state-space model

2021-09-23 · Jeremy H. M. Wong, Yifan Gong

Speakers may move around while diarisation is being performed. When a microphone array is used, the instantaneous locations of where the sounds originated from can be estimated, and previous investigations have shown that such information can be complementary to speaker embeddings in the diarisation task. However, these approaches often assume that speakers are fairly stationary throughout a meeting. This paper relaxes this assumption, by proposing to explicitly track the movements of speakers while jointly performing diarisation within a unified model. A state-space model is proposed, where the hidden state expresses the identity of the current active speaker and the predicted locations of all speakers. The model is implemented as a particle filter. Experiments on a Microsoft rich meeting transcription task show that the proposed joint location tracking and diarisation approach is able to perform comparably with other methods that use location information.

📄 PDF Abstract BibTeX arXiv:2109.11140

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Diarisation using location tracking with agglomerative clustering

2021-09-22 · Jeremy H. M. Wong, Igor Abramovski, Xiong Xiao, Yifan Gong

Previous works have shown that spatial location information can be complementary to speaker embeddings for a speaker diarisation task. However, the models used often assume that speakers are fairly stationary throughout …

Clustering

Detecting agreement in multi-party dialogue: evaluating speaker diarisation versus a procedural baseline to enhance user engagement

2023-11-06 · Angus Addlesee, Daniel Denley, Andy Edmondson, Nancie Gunson 외

Conversational agents participating in multi-party interactions face significant challenges in dialogue state tracking, since the identity of the speaker adds significant contextual meaning. It is common to utilise diari…

Dialogue State Tracking

In search of strong embedding extractors for speaker diarisation

2022-10-26 · Jee-weon Jung, Hee-Soo Heo, Bong-Jin Lee, Jaesung Huh 외

Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisatio…

Data AugmentationSpeaker Verification

Adapting Speaker Embeddings for Speaker Diarisation

2021-04-07 · Youngki Kwon, Jee-weon Jung, Hee-Soo Heo, You Jin Kim 외

The goal of this paper is to adapt speaker embeddings for solving the problem of speaker diarisation. The quality of speaker embeddings is paramount to the performance of speaker diarisation systems. Despite this, prior …

ClusteringDimensionality ReductionSpeaker Verification

Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription

2022-07-08 · Xianrui Zheng, Chao Zhang, Philip C. Woodland

Self-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diarisation and speech recognition using a s…

Action DetectionActivity DetectionSelf-Supervised Learningspeech-recognition+1