paper-with-me

홈 › Papers

Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task

2025-09-04 · JungHo Jung, Junhyun Lee arxiv

End-to-end speech-to-text translation typically suffers from the scarcity of paired speech-text data. One way to overcome this shortcoming is to utilize the bitext data from the Machine Translation (MT) task and perform Multi-Task Learning (MTL). In this paper, we formulate MTL from a regularization perspective and explore how sequences can be regularized within and across modalities. By thoroughly investigating the effect of consistency regularization (different modality) and R-drop (same modality), we show how they respectively contribute to the total regularization. We also demonstrate that the coefficient of MT loss serves as another source of regularization in the MTL setting. With these three sources of regularization, we introduce the optimal regularization contour in the high-dimensional space, called the regularization horizon. Experiments show that tuning the hyperparameters within the regularization horizon achieves near state-of-the-art performance on the MuST-C dataset.

📄 PDF Abstract BibTeX arXiv:2509.09701

Code (0)

등록된 구현이 없습니다.

Tasks

Speech-to-Text TranslationMachine TranslationMulti-Task Learning

Similar Papers 제목 키워드 기반

Optimal Transport Regularization for Speech Text Alignment in Spoken Language Models

2025-08-11 · Wenze Xu, Chun Wang, Jiazhen Yu, Sheng Chen 외 arxiv

Spoken Language Models (SLMs), which extend Large Language Models (LLMs) to perceive speech inputs, have gained increasing attention for their potential to advance speech understanding tasks. However, despite recent prog…

Time Regularization in Optimal Time Variable Learning

2023-06-28 · Evelyn Herberg, Roland Herzog, Frederik Köhne

Recently, optimal time variable learning in deep neural networks (DNNs) was introduced in arXiv:2204.08528. In this manuscript we extend the concept by introducing a regularization term that directly relates to the time …

Modality Adaption or Regularization? A Case Study on End-to-End Speech Translation

2023-06-13 · Yuchen Han, Chen Xu, Tong Xiao, Jingbo Zhu

Pre-training and fine-tuning is a paradigm for alleviating the data scarcity problem in end-to-end speech translation (E2E ST). The commonplace "modality gap" between speech and text data often leads to inconsistent inpu…

Population Based Training for Data Augmentation and Regularization in Speech Recognition

2020-10-08 · Daniel Haziza, Jérémy Rapin, Gabriel Synnaeve

Varying data augmentation policies and regularization over the course of optimization has led to performance improvements over using fixed values. We show that population based training is a useful tool to continuously s…

Data Augmentationspeech-recognitionSpeech Recognition

AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition

2024-03-18 · SooHwan Eom, Eunseop Yoon, Hee Suk Yoon, Chanwoo Kim 외

In Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a r…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition