paper-with-me

홈 › Papers

MIRNet: Learning multiple identities representations in overlapped speech

2020-08-04 · Hyewon Han, Soo-Whan Chung, Hong-Goo Kang

Many approaches can derive information about a single speaker's identity from the speech by learning to recognize consistent characteristics of acoustic parameters. However, it is challenging to determine identity information when there are multiple concurrent speakers in a given signal. In this paper, we propose a novel deep speaker representation strategy that can reliably extract multiple speaker identities from an overlapped speech. We design a network that can extract a high-level embedding that contains information about each speaker's identity from a given mixture. Unlike conventional approaches that need reference acoustic features for training, our proposed algorithm only requires the speaker identity labels of the overlapped speech segments. We demonstrate the effectiveness and usefulness of our algorithm in a speaker verification task and a speech separation system conditioned on the target speaker embeddings obtained through the proposed method.

📄 PDF Abstract BibTeX arXiv:2008.01698

Code (0)

등록된 구현이 없습니다.

Tasks

Rgb-T TrackingSpeaker VerificationSpeech Separation

Similar Papers 제목 키워드 기반

UICE-MIRNet guided image enhancement for underwater object detection

2024-09-24 · Nature 2024 9 · Pratima Sarkar1, 3, Sourav De2, Sandeep Gurung1 & Prasenjit Dey4

Underwater object detection is a crucial aspect of monitoring the aquaculture resources to preserve the marine ecosystem. In most cases, Low-light and scattered lighting conditions create challenges for computer vision…

feature selectionImage EnhancementLow-Light Image EnhancementObject+2

Serialized Output Training for End-to-End Overlapped Speech Recognition

2020-03-28 · Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng 외

This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach. Instead of having multiple output layers as wi…

Decoderspeech-recognitionSpeech Recognition

Three-class Overlapped Speech Detection using a Convolutional Recurrent Neural Network

2021-04-07 · Jee-weon Jung, Hee-Soo Heo, Youngki Kwon, Joon Son Chung 외

In this work, we propose an overlapped speech detection system trained as a three-class classifier. Unlike conventional systems that perform binary classification as to whether or not a frame contains overlapped speech, …

Binary Classificationspeaker-diarizationSpeaker Diarization

Continuous speech separation: dataset and analysis

2020-01-30 · Zhuo Chen, Takuya Yoshioka, Liang Lu, Tianyan Zhou 외

This paper describes a dataset and protocols for evaluating continuous speech separation algorithms. Most prior studies on speech separation use pre-segmented signals of artificially mixed speech utterances which are mos…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers

2020-06-19 · Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng 외

We propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped speech. Our model is built on serialized…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderSpeaker Identification+2