paper-with-me

Papers

3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

2023-06-27 · Siqi Zheng, Luyao Cheng, Yafeng Chen, Hui Wang, Qian Chen

Disentangling uncorrelated information in speech utterances is a crucial research topic within speech community. Different speech-related tasks focus on extracting distinct speech representations while minimizing the affects of other uncorrelated information. We present a large-scale speech corpus to facilitate the research of speech representation disentanglement. 3D-Speaker contains over 10,000 speakers, each of whom are simultaneously recorded by multiple Devices, locating at different Distances, and some speakers are speaking multiple Dialects. The controlled combinations of multi-dimensional audio data yield a matrix of a diverse blend of speech representation entanglement, thereby motivating intriguing methods to untangle them. The multi-domain nature of 3D-Speaker also makes it a suitable resource to evaluate large universal speech models and experiment methods of out-of-domain learning and self-supervised learning. https://3dspeaker.github.io/

📄 PDF Abstract BibTeX arXiv:2306.15354

Code (2)

alibaba-damo-academy/3D-Speaker 공식 구현 pytorch
modelscope/3D-Speaker pytorch

Tasks

DisentanglementSelf-Supervised Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

End-to-end Alexa Device Arbitration

2021-12-08 · Jarred Barber, Yifeng Fan, Tao Zhang

We introduce a variant of the speaker localization problem, which we call device arbitration. In the device arbitration problem, a user utters a keyword that is detected by multiple distributed microphone arrays (smart h…

On-Device Speaker Anonymization of Acoustic Embeddings for ASR based onFlexible Location Gradient Reversal Layer

2023-07-25 · Md Asif Jalal, Pablo Peso Parada, Jisi Zhang, Karthikeyan Saravanan 외

Smart devices serviced by large-scale AI models necessitates user data transfer to the cloud for inference. For speech applications, this means transferring private user information, e.g., speaker identity. Our paper pro…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker anonymizationSpeaker Recognition+2

Discussion on domain generalization in the cross-device speaker verification system

2021-10-01 · ROCLING 2021 10 · Wei-Ting Lin, Yu-jia Zhang, Chia-Ping Chen, Chung-Li Lu 외

In this paper, we use domain generalization to improve the performance of the cross-device speaker verification system. Based on a trainable speaker verification system, we use domain generalization algorithms to fine-tu…

Domain GeneralizationSpeaker Verification

Keyword Spotting for Hearing Assistive Devices Robust to External Speakers

2019-06-22 · Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen

Keyword spotting (KWS) is experiencing an upswing due to the pervasiveness of small electronic devices that allow interaction with them via speech. Often, KWS systems are speaker-independent, which means that any person …

Keyword SpottingMulti-Task Learning

Highly Efficient Real-Time Streaming and Fully On-Device Speaker Diarization with Multi-Stage Clustering

2022-10-25 · Quan Wang, Yiling Huang, Han Lu, Guanlong Zhao 외

While recent research advances in speaker diarization mostly focus on improving the quality of diarization results, there is also an increasing interest in improving the efficiency of diarization systems. In this paper, …

ClusteringCPUspeaker-diarizationSpeaker Diarization