paper-with-me

Papers

A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement

2024-03-03 · Ravi Shankar, Ke Tan, Buye Xu, Anurag Kumar

Self-supervised learned models have been found to be very effective for certain speech tasks such as automatic speech recognition, speaker identification, keyword spotting and others. While the features are undeniably useful in speech recognition and associated tasks, their utility in speech enhancement systems is yet to be firmly established, and perhaps not properly understood. In this paper, we investigate the uses of SSL representations for single-channel speech enhancement in challenging conditions and find that they add very little value for the enhancement task. Our constraints are designed around on-device real-time speech enhancement -- model is causal, the compute footprint is small. Additionally, we focus on low SNR conditions where such models struggle to provide good enhancement. In order to systematically examine how SSL representations impact performance of such enhancement models, we propose a variety of techniques to utilize these embeddings which include different forms of knowledge-distillation and pre-training.

📄 PDF Abstract BibTeX arXiv:2403.01369

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionKeyword SpottingKnowledge DistillationSpeaker IdentificationSpeech Enhancementspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Device Passport: Enabling Spatio-Temporal Pretrained Models to Generalize Across Input Layouts

2026-06-30 · Geeling Chau, Ran Liu, Juri Minxha, Wenhui Cui 외 arxiv

New device layouts pose a challenging modeling problem due to the lack of large datasets for each specific layout. Biosignal foundation models offer a plausible solution if they are able to generalize to new layouts effe…

Quantitative Evidence on Overlooked Aspects of Enrollment Speaker Embeddings for Target Speaker Separation

2022-10-23 · Xiaoyu Liu, Xu Li, Joan Serrà

Single channel target speaker separation (TSS) aims at extracting a speaker's voice from a mixture of multiple talkers given an enrollment utterance of that speaker. A typical deep learning TSS framework consists of an u…

Speaker IdentificationSpeaker Separation

Deep Learning for Joint Source-Channel Coding of Text

2018-02-19 · Nariman Farsad, Milind Rao, Andrea Goldsmith

We consider the problem of joint source and channel coding of structured data such as natural language over a noisy channel. The typical approach to this problem in both theory and practice involves performing source cod…

DecoderDeep Learning

A Closer Look at Few-Shot 3D Point Cloud Classification

2023-03-31 · Chuangguan Ye, Hongyuan Zhu, Bo Zhang, Tao Chen

In recent years, research on few-shot learning (FSL) has been fast-growing in the 2D image domain due to the less requirement for labeled training data and greater generalization for novel classes. However, its applicati…

3D Point Cloud ClassificationFew-Shot 3D Point Cloud ClassificationFew-Shot LearningPoint Cloud Classification

A Closer Look on Unsupervised Cross-lingual Word Embeddings Mapping

2020-05-01 · LREC 2020 5 · Kamil Pluci{\'n}ski, Mateusz Lango, Micha{\l} Zimniewicz

In this work, we study the unsupervised cross-lingual word embeddings mapping method presented by Artetxe et al. (2018). First, wesuccessfully reproduced the experiments performed in the original work, finding only minor…

Cross-Lingual Word EmbeddingsWord Embeddings