paper-with-me

홈 › Papers

Cross-Modal Knowledge Distillation for Speech Large Language Models

2025-09-18 · Enzhi Wang, Qicheng Li, Zhiyuan Tang, Yuhang Jia arxiv

In this work, we present the first systematic evaluation of catastrophic forgetting and modality inequivalence in speech large language models, showing that introducing speech capabilities can degrade knowledge and reasoning even when inputs remain textual, and performance further decreases with spoken queries. To address these challenges, we propose a cross-modal knowledge distillation framework that leverages both text-to-text and speech-to-text channels to transfer knowledge from a text-based teacher model to a speech LLM. Extensive experiments on dialogue and audio understanding tasks validate the effectiveness of our approach in preserving textual knowledge, improving cross-modal alignment, and enhancing reasoning in speech-based interactions.

📄 PDF Abstract BibTeX arXiv:2509.14930

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

Knowledge distillation from language model to acoustic model: a hierarchical multi-task learning approach

2021-10-20 · Mun-Hak Lee, Joon-Hyuk Chang

The remarkable performance of the pre-trained language model (LM) using self-supervised learning has led to a major paradigm shift in the study of natural language processing. In line with these changes, leveraging the p…

Knowledge DistillationLanguage ModelingLanguage Modellingmodel+4

Knowledge Transfer from Pre-trained Language Models to Cif-based Speech Recognizers via Hierarchical Distillation

2023-01-30 · Minglun Han, Feilong Chen, Jing Shi, Shuang Xu 외

Large-scale pre-trained language models (PLMs) have shown great potential in natural language processing tasks. Leveraging the capabilities of PLMs to enhance automatic speech recognition (ASR) systems has also emerged a…

Automatic Speech RecognitionKnowledge DistillationLanguage Modellingspeech-recognition+2

Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models

2025-02-09 · Jing-Xuan Zhang, Genshun Wan, Jianqing Gao, Zhen-Hua Ling

Audio-visual representation learning is crucial for advancing multimodal speech processing tasks, such as lipreading and audio-visual speech recognition. Recently, speech foundation models (SFMs) have shown remarkable ge…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge Distillation+5

Closing the Gap Between Text and Speech Understanding in LLMs

2025-10-15 · Santiago Cuervo, Skyler Seto, Maureen de Seyssel, Richard He Bai 외 arxiv

Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counterparts--and even cascaded pipelines--on …

Speech Synthesis

Adaptive Knowledge Distillation between Text and Speech Pre-trained Models

2023-03-07 · Jinjie Ni, Yukun Ma, Wen Wang, Qian Chen 외

Learning on a massive amount of speech corpus leads to the recent success of many self-supervised speech models. With knowledge distillation, these models may also benefit from the knowledge encoded by language models th…

Knowledge DistillationSpoken Language Understanding