paper-with-me

Papers

Front-End Adapter: Adapting Front-End Input of Speech based Self-Supervised Learning for Speech Recognition

2023-02-18 · Xie Chen, Ziyang Ma, Changli Tang, Yujin Wang, Zhisheng Zheng

Recent years have witnessed a boom in self-supervised learning (SSL) in various areas including speech processing. Speech based SSL models present promising performance in a range of speech related tasks. However, the training of SSL models is computationally expensive and a common practice is to fine-tune a released SSL model on the specific task. It is essential to use consistent front-end input during pre-training and fine-tuning. This consistency may introduce potential issues when the optimal front-end is not the same as that used in pre-training. In this paper, we propose a simple but effective front-end adapter to address this front-end discrepancy. By minimizing the distance between the outputs of different front-ends, the filterbank feature (Fbank) can be compatible with SSL models which are pre-trained with waveform. The experiment results demonstrate the effectiveness of our proposed front-end adapter on several popular SSL models for the speech recognition task.

📄 PDF Abstract BibTeX arXiv:2302.09331

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learningspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy Conditions

2024-04-10 · Hanyu Meng, Vidhyasaharan Sethu, Eliathamby Ambikairajah

There is increasing interest in the use of the LEArnable Front-end (LEAF) in a variety of speech processing systems. However, there is a dearth of analyses of what is actually learnt and the relative importance of traini…

Emotion RecognitionKeyword SpottingLanguage Identification

Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters

2024-01-10 · Kenichi Fujita, Hiroshi Sato, Takanori Ashihara, Hiroki Kanagawa 외

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. H…

Self-Supervised LearningSpeech EnhancementSpeech Synthesistext-to-speech+2

Exploration of Adapter for Noise Robust Automatic Speech Recognition

2024-02-28 · Hao Shi, Tatsuya Kawahara

Adapting an automatic speech recognition (ASR) system to unseen noise environments is crucial. Integrating adapters into neural networks has emerged as a potent technique for transfer learning. This study thoroughly inve…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+2

Speaker Characterization by means of Attention Pooling

2024-05-07 · Federico Costa, Miquel India, Javier Hernando

State-of-the-art Deep Learning systems for speaker verification are commonly based on speaker embedding extractors. These architectures are usually composed of a feature extractor front-end together with a pooling layer …

Emotion RecognitionSpeaker RecognitionSpeaker Verification

A Monaural Speech Enhancement Method for Robust Small-Footprint Keyword Spotting

2019-06-20 · Yue Gu, Zhihao Du, HUI ZHANG, Xueliang Zhang

Robustness against noise is critical for keyword spotting (KWS) in real-world environments. To improve the robustness, a speech enhancement front-end is involved. Instead of treating the speech enhancement as a separated…

Keyword SpottingSmall-Footprint Keyword SpottingSpeech Enhancement