paper-with-me

Papers

Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning

2025-06-06 · Yangui Fang, Jing Peng, Xu Li, Yu Xi, Chengwei Zhang, Guohui Zhong, Kai Yu

Recent advances in automatic speech recognition (ASR) have combined speech encoders with large language models (LLMs) through projection, forming Speech LLMs with strong performance. However, adapting them to new domains remains challenging, especially in low-resource settings where paired speech-text data is scarce. We propose a text-only fine-tuning strategy for Speech LLMs using unpaired target-domain text without requiring additional audio. To preserve speech-text alignment, we introduce a real-time evaluation mechanism during fine-tuning. This enables effective domain adaptation while maintaining source-domain performance. Experiments on LibriSpeech, SlideSpeech, and Medical datasets show that our method achieves competitive recognition performance, with minimal degradation compared to full audio-text fine-tuning. It also improves generalization to new domains without catastrophic forgetting, highlighting the potential of text-only fine-tuning for low-resource domain adaptation of ASR.

📄 PDF Abstract BibTeX arXiv:2506.05671

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning

2026-01-26 · Haolong Zheng, Siyin Wang, Zengrui Jin, Mark Hasegawa-Johnson arxiv

Auditory Large Language Models (LLMs) have demonstrated strong performance across a wide range of speech and audio understanding tasks. Nevertheless, they often struggle when applied to low-resource tasks. In case in-dom…

TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs

2026-04-09 · Jing Peng, Chenghao Wang, Yi Yang, Lirong Qian 외 arxiv

Speech LLM post-training increasingly relies on efficient cross-modal alignment and robust low-resource adaptation, yet collecting large-scale audio-text pairs remains costly. Text-only alignment methods such as TASU red…

Text-only Domain Adaptation using Unified Speech-Text Representation in Transducer

2023-06-07 · Lu Huang, Boyu Li, Jun Zhang, Lu Lu 외

Domain adaptation using text-only corpus is challenging in end-to-end(E2E) speech recognition. Adaptation by synthesizing audio from text through TTS is resource-consuming. We present a method to learn Unified Speech-Tex…

Domain AdaptationLanguage ModelingLanguage Modellingspeech-recognition+1

SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR

2024-06-15 · Natarajan Balaji Shankar, Ruchao Fan, Abeer Alwan

Recently, speech foundation models have gained popularity due to their superiority in finetuning downstream ASR tasks. However, models finetuned on certain domains, such as LibriSpeech (adult read speech), behave poorly …

Domain Adaptation

SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition

2024-09-16 · Ming-Hao Hsu, Hung-Yi Lee

Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due to limited training data and insufficien…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationIn-Context Learning+4