paper-with-me

홈 › Papers

Adapting Text LLMs to Speech via Multimodal Depth Up-Scaling

2026-04-01 · Kazuki Yano, Jun Suzuki, Shinji Watanabe arxiv

Adapting pre-trained text Large Language Models (LLMs) into Speech Language Models (Speech LMs) via continual pretraining on speech data is promising, but often degrades the original text capabilities. We propose Multimodal Depth Upscaling, an extension of an emerging strategy in continual LLM pre-training, where new transformer layers are inserted into a frozen text LLM and only the added layers are trained on speech data. Experiments with SmolLM2-360M and SmolLM2-1.7B on 48k hours of English Automatic Speech Recognition (ASR) data show that depth up-scaling achieves ASR comparable to full fine-tuning while causing far less text degradation than both full fine-tuning and Low-Rank Adaptation (LoRA). We further show that incorporating E-Branchformer, an architecture designed for speech recognition, as the inserted layers achieves ASR that matches or surpasses full fine-tuning on the larger model while reducing text degradation by over 75% with 60% fewer trainable parameters.

📄 PDF Abstract BibTeX arXiv:2604.00489

Code (0)

등록된 구현이 없습니다.

Tasks

Continual PretrainingSpeech Recognition

Similar Papers 제목 키워드 기반

MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning

2026-01-26 · Haolong Zheng, Siyin Wang, Zengrui Jin, Mark Hasegawa-Johnson arxiv

Auditory Large Language Models (LLMs) have demonstrated strong performance across a wide range of speech and audio understanding tasks. Nevertheless, they often struggle when applied to low-resource tasks. In case in-dom…

ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs

2025-05-26 · Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli, Cem Subakan

Large Language Models (LLMs) are widely used in Spoken Language Understanding (SLU). Recent SLU models process audio directly by adapting speech input into LLMs for better multimodal learning. A key consideration for the…

cross-modal alignmentEmotion RecognitionQuestion AnsweringSpoken Language Understanding

A Survey on Speech Large Language Models

2024-10-24 · Jing Peng, Yucheng Wang, Yangui Fang, Yu Xi 외

Large Language Models (LLMs) exhibit strong contextual understanding and remarkable multitask performance. As a result, researchers have been actively exploring the integration of LLMs into the domain of speech understan…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognition+6

Tuning Large language model for End-to-end Speech Translation

2023-10-03 · Hao Zhang, Nianwen Si, Yaqi Chen, Wenlin Zhang 외

With the emergence of large language models (LLMs), multimodal models based on LLMs have demonstrated significant potential. Models such as LLaSM, X-LLM, and SpeechGPT exhibit an impressive ability to comprehend and gene…

de-enfr-enLanguage ModelingLanguage Modelling+3

Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition

2023-07-17 · Shaoshi Ling, Yuxuan Hu, Shuangbei Qian, Guoli Ye 외

Most end-to-end (E2E) speech recognition models are composed of encoder and decoder blocks that perform acoustic and language modeling functions. Pretrained large language models (LLMs) have the potential to improve the …

DecoderLanguage ModelingLanguage ModellingLarge Language Model+2