paper-with-me

홈 › Papers

Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech Translation

2019-09-17 · Chengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang, Ming Zhou

End-to-end speech translation, a hot topic in recent years, aims to translate a segment of audio into a specific language with an end-to-end model. Conventional approaches employ multi-task learning and pre-training methods for this task, but they suffer from the huge gap between pre-training and fine-tuning. To address these issues, we propose a Tandem Connectionist Encoding Network (TCEN) which bridges the gap by reusing all subnets in fine-tuning, keeping the roles of subnets consistent, and pre-training the attention module. Furthermore, we propose two simple but effective methods to guarantee the speech encoder outputs and the MT encoder inputs are consistent in terms of semantic representation and sequence length. Experimental results show that our model outperforms baselines 2.2 BLEU on a large benchmark dataset.

📄 PDF Abstract BibTeX arXiv:1909.07575

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task LearningTranslation

Similar Papers 제목 키워드 기반

Multilingual Auxiliary Tasks Training: Bridging the Gap between Languages for Zero-Shot Transfer of Hate Speech Detection Models

2022-10-24 · Syrielle Montariol, Arij Riabi, Djamé Seddah

Zero-shot cross-lingual transfer learning has been shown to be highly challenging for tasks involving a lot of linguistic specificities or when a cultural gap is present between languages, such as in hate speech detectio…

Cross-Lingual TransferHate Speech Detectionnamed-entity-recognitionNamed Entity Recognition+4

Post-training for Deepfake Speech Detection

2025-06-26 · Wanying Ge, Xin Wang, Xuechen Liu, Junichi Yamagishi

We introduce a post-training approach that adapts self-supervised learning (SSL) models for deepfake speech detection by bridging the gap between general pre-training and domain-specific fine-tuning. We present AntiDeepf…

Face SwappingSelf-Supervised Learning

When More is not Necessary Better: Multilingual Auxiliary Tasks for Zero-Shot Cross-Lingual Transfer of Hate Speech Detection Models

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Zero-shot cross-lingual transfer learning has been shown to be highly challenging for tasks involving a lot of linguistic specificities or when a cultural gap is present between languages, such as in hate speech detectio…

Cross-Lingual TransferHate Speech DetectionLanguage ModelingLanguage Modelling+6

Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR

2026-01-04 · Yuxiang Mei, Dongxing Xu, Jiaen Liang, Yanhua Long arxiv

The INTERSPEECH 2025 Challenge on Multilingual Conversational Speech Language Models (MLC-SLM) promotes multilingual conversational ASR with large language models (LLMs). Our previous SHNU-mASR system adopted a competiti…

SpeechCLIP: Integrating Speech with Pre-Trained Vision and Language Model

2022-10-03 · Yi-Jen Shih, Hsuan-Fu Wang, Heng-Jui Chang, Layne Berry 외

Data-driven speech processing models usually perform well with a large amount of text supervision, but collecting transcribed speech data is costly. Therefore, we propose SpeechCLIP, a novel framework bridging speech and…

Language ModelingLanguage ModellingRetrievalText Retrieval