paper-with-me

홈 › Papers

RedApt: An Adaptor for wav2vec 2 Encoding \\ Faster and Smaller Speech Translation without Quality Compromise

2022-10-16 · Jinming Zhao, Hao Yang, Gholamreza Haffari, Ehsan Shareghi

Pre-trained speech Transformers in speech translation (ST) have facilitated state-of-the-art (SotA) results; yet, using such encoders is computationally expensive. To improve this, we present a novel Reducer Adaptor block, RedApt, that could be seamlessly integrated within any Transformer-based speech encoding architecture. Integrating the pretrained wav2vec 2 speech encoder with RedAptbrings 41% speedup, 33% memory reduction with 24% fewer FLOPs at inference. To our positive surprise, our ST model with RedApt outperforms the SotA architecture by an average of 0.68 BLEU score on 8 language pairs from Must-C.

📄 PDF Abstract BibTeX arXiv:2210.08475

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor

2025-05-26 · Seokgi Lee, Jungjun Kim

We present the gradual style adaptor TTS (GSA-TTS) with a novel style encoder that gradually encodes speaking styles from an acoustic reference for zero-shot speech synthesis. GSA first captures the local style of each s…

Speech Synthesis

Universal Adaptor: Converting Mel-Spectrograms Between Different Configurations for Speech Synthesis

2022-04-01 · Fan-Lin Wang, Po-chun Hsu, Da-Rong Liu, Hung-Yi Lee

Most recent speech synthesis systems are composed of a synthesizer and a vocoder. However, the existing synthesizers and vocoders can only be matched to acoustic features extracted with a specific configuration. Hence, w…

Speech SynthesisVoice Conversion

Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders

2021-05-12 · ACL 2021 5 · Chen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang 외

Encoder pre-training is promising in end-to-end Speech Translation (ST), given the fact that speech-to-translation data is scarce. But ST encoders are not simple instances of Automatic Speech Recognition (ASR) or Machine…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge DistillationMachine Translation+3

GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech

2022-05-15 · Rongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui 외

Style transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following chal…

Speech SynthesisStyle Transfertext-to-speechText to Speech+1

End-to-End Simultaneous Dysarthric Speech Reconstruction with Frame-Level Adaptor and Multiple Wait-k Knowledge Distillation

2026-03-02 · Minghui Wu, Haitao Tang, Jiahuan Fan, Ruizhi Liao 외 arxiv

Dysarthric speech reconstruction (DSR) typically employs a cascaded system that combines automatic speech recognition (ASR) and sentence-level text-to-speech (TTS) to convert dysarthric speech into normally-prosodied spe…

Knowledge DistillationSpeech Recognition