paper-with-me

홈 › Papers

Disentangling the Impacts of Language and Channel Variability on Speech Separation Networks

2022-03-30 · Fan-Lin Wang, Hung-Shin Lee, Yu Tsao, Hsin-Min Wang

Because the performance of speech separation is excellent for speech in which two speakers completely overlap, research attention has been shifted to dealing with more realistic scenarios. However, domain mismatch between training/test situations due to factors, such as speaker, content, channel, and environment, remains a severe problem for speech separation. Speaker and environment mismatches have been studied in the existing literature. Nevertheless, there are few studies on speech content and channel mismatches. Moreover, the impacts of language and channel in these studies are mostly tangled. In this study, we create several datasets for various experiments. The results show that the impacts of different languages are small enough to be ignored compared to the impacts of different channels. In our experiments, training on data recorded by Android phones leads to the best generalizability. Moreover, we provide a new solution for channel mismatch by evaluating projection, where the channel similarity can be measured and used to effectively select additional training data to improve the performance of in-the-wild test data.

📄 PDF Abstract BibTeX arXiv:2203.16040

Code (1)

sinica-slam/cospro-mix 공식 구현

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Subject Disentanglement Neural Network for Speech Envelope Reconstruction from EEG

2025-01-15 · Li Zhang, Jiyao Liu

Reconstructing speech envelopes from EEG signals is essential for exploring neural mechanisms underlying speech perception. Yet, EEG variability across subjects and physiological artifacts complicate accurate reconstruct…

DisentanglementEEGEeg Decoding

UR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021

2021-07-26 · Xinhui Chen, You Zhang, Ge Zhu, Zhiyao Duan

In this paper, we present UR-AIR system submission to the logical access (LA) and the speech deepfake (DF) tracks of the ASVspoof 2021 Challenge. The LA and DF tasks focus on synthetic speech detection (SSD), i.e. detect…

Audio CompressionFace SwappingSynthetic Speech Detectiontext-to-speech+2

Joint Probabilistic Linear Discriminant Analysis

2017-04-07 · Luciana Ferrer

Standard probabilistic linear discriminant analysis (PLDA) for speaker recognition assumes that the sample's features (usually, i-vectors) are given by a sum of three terms: a term that depends on the speaker identity, a…

Speaker Recognition

Reducing Language confusion for Code-switching Speech Recognition with Token-level Language Diarization

2022-10-26 · Hexin Liu, HaiHua Xu, Leibny Paola Garcia, Andy W. H. Khong 외

Code-switching (CS) refers to the phenomenon that languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). This paper aims to address language confusion for improvin…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Stochastic density effects on adult fish survival and implications for population fluctuations

2016-01-22

The degree to which population fluctuations arise from variable adult survival relative to variable recruitment has been debated widely for marine organisms. Disentangling these effects remains challenging because data g…

Time SeriesTime Series Analysis