paper-with-me

Papers

Transferable speech-to-text large language model alignment module

2024-06-19 · Boyong Wu, Chao Yan, Haoran Pu

By leveraging the power of Large Language Models(LLMs) and speech foundation models, state of the art speech-text bimodal works can achieve challenging tasks like spoken translation(ST) and question answering(SQA) altogether with much simpler architectures. In this paper, we utilize the capability of Whisper encoder and pre-trained Yi-6B. Empirical results reveal that modal alignment can be achieved with one layer module and hundred hours of speech-text multitask corpus. We further swap the Yi-6B with human preferences aligned version of Yi-6B-Chat during inference, and discover that the alignment capability is applicable as well. In addition, the alignment subspace revealed by singular value decomposition(SVD) also implies linear alignment subspace is sparse, which leaves the possibility to concatenate other features like voice-print or video to expand modality.

📄 PDF Abstract BibTeX arXiv:2406.13357

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelQuestion AnsweringSpeech-to-Text

Similar Papers 제목 키워드 기반

Representation Potentials of Foundation Models for Multimodal Alignment: A Survey

2025-10-05 · Jianglin Lu, Hailing Wang, Yi Xu, Yizhou Wang 외 arxiv

Foundation models learn highly transferable representations through large-scale pretraining on diverse data. An increasing body of research indicates that these representations exhibit a remarkable degree of similarity a…

PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs

2025-09-24 · Pei Zhang, Andong Chen, Xi Chen, Baosong Yang 외 arxiv

Large language models (LLMs) have expanded from text to speech, giving rise to Speech Large Models (SLMs) that support recognition, translation, and synthesis. A key challenge is aligning speech and text representations,…

Text to Speech

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

2025-06-16 · Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou 외

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatena…

Large Language Modelmultimodal interaction

StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis

2022-05-30 · Yinghao Aaron Li, Cong Han, Nima Mesgarani

Text-to-Speech (TTS) has recently seen great progress in synthesizing high-quality speech owing to the rapid development of parallel TTS systems, but producing speech with naturalistic prosodic variations, speaking style…

Data AugmentationSelf-Supervised LearningSpeech Synthesistext-to-speech+2

BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

2023-09-02 · Chen Wang, Minpeng Liao, Zhongqiang Huang, Jinliang Lu 외

The emergence of large language models (LLMs) has sparked significant interest in extending their remarkable language capabilities to speech. However, modality alignment between speech and text still remains an open prob…

speech-recognitionSpeech RecognitionSpoken Language Understanding