paper-with-me

홈 › Papers

Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts

2025-08-05 · Hojun Jin, Eunsoo Hong, Ziwon Hyung, Sungjun Lim, Seungjin Lee, Keunseok Cho arxiv

Hard-parameter sharing is a common strategy to train a single model jointly across diverse tasks. However, this often leads to task interference, impeding overall model performance. To address the issue, we propose a simple yet effective Supervised Mixture of Experts (S-MoE). Unlike traditional Mixture of Experts models, S-MoE eliminates the need for training gating functions by utilizing special guiding tokens to route each task to its designated expert. By assigning each task to a separate feedforward network, S-MoE overcomes the limitations of hard-parameter sharing. We further apply S-MoE to a speech-to-text model, enabling the model to process mixed-bandwidth input while jointly performing automatic speech recognition (ASR) and speech translation (ST). Experimental results demonstrate the effectiveness of the proposed S-MoE, achieving a 6.35% relative improvement in Word Error Rate (WER) when applied to both the encoder and decoder.

📄 PDF Abstract BibTeX arXiv:2508.10009

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing

2023-09-27 · Brian Yan, Xuankai Chang, Antonios Anastasopoulos, Yuya Fujita 외

Recent works in end-to-end speech-to-text translation (ST) have proposed multi-tasking methods with soft parameter sharing which leverage machine translation (MT) data via secondary encoders that map text inputs to an ev…

DecoderMachine TranslationSpeech-to-TextSpeech-to-Text Translation+2

iCompass at Arabic Hate Speech 2022: Detect Hate Speech Using QRNN and Transformers

2022-06-01 · OSACT (LREC) 2022 6 · Mohamed Aziz Bennessir, Malek Rhouma, Hatem Haddad, Chayma Fourati

This paper provides a detailed overview of the system we submitted as part of the OSACT2022 Shared Tasks on Fine-Grained Hate Speech Detection on Arabic Twitter, its outcome, and limitations. Our submission is accomplish…

Hate Speech Detection

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

2025-12-16 · Juze Zhang, Changan Chen, Xin Chen, Heng Yu 외 arxiv

Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human behavior as a translation task co-speech gesture or text-to-motion that …

Learning Sparse Sharing Architectures for Multiple Tasks

2019-11-12 · Tianxiang Sun, Yunfan Shao, Xiaonan Li, PengFei Liu 외

Most existing deep multi-task learning models are based on parameter sharing, such as hard sharing, hierarchical sharing, and soft sharing. How choosing a suitable sharing mechanism depends on the relations among the tas…

Multi-Task Learning

Risk of re-identification for shared clinical speech recordings

2022-10-18 · Daniela A. Wiepert, Bradley A. Malin, Joseph R. Duffy, Rene L. Utianski 외

Large, curated datasets are required to leverage speech-based tools in healthcare. These are costly to produce, resulting in increased interest in data sharing. As speech can potentially identify speakers (i.e., voicepri…

Speaker Recognition