DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
Recent speech language models (SLMs) typically incorporate pre-trained speech models to extend the capabilities from large language models (LLMs). In this paper, we propose a Descriptive Speech-Text Alignment approach that leverages speech captioning to bridge the gap between speech and text modalities, enabling SLMs to interpret and generate comprehensive natural language descriptions, thereby facilitating the capability to understand both linguistic and non-linguistic features in speech. Enhanced with the proposed approach, our model demonstrates superior performance on the Dynamic-SUPERB benchmark, particularly in generalizing to unseen tasks. Moreover, we discover that the aligned model exhibits a zero-shot instruction-following capability without explicit speech instruction tuning. These findings highlight the potential to reshape instruction-following SLMs by incorporating rich, descriptive speech captions.
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveInstruction FollowingSimilar Papers 제목 키워드 기반
SpeakGer: A meta-data enriched speech corpus of German state and federal parliaments
The application of natural language processing on political texts as well as speeches has become increasingly relevant in political sciences due to the ability to analyze large text corpora which cannot be read by a sing…
DescriptiveSentiment AnalysisDeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following, without requiring task-specific audio instruction-tuning. Recent LALMs t…
cross-modal alignmentInstruction FollowingLanguage ModelingLanguage Modelling+1Multilingual and Explainable Text Detoxification with Parallel Corpora
Even with various regulations in place across countries and social media platforms (Government of India, 2021; European Parliament and Council of the European Union, 2022, digital abusive speech remains a significant iss…
DescriptiveStyle TransferText Style TransferSemantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
Speech Emotion Recognition (SER) is essential for improving human-computer interaction, yet its accuracy remains constrained by the complexity of emotional nuances in speech. In this study, we distinguish between descrip…
Speech Emotion RecognitionEnhancing Descriptive Image Captioning with Natural Language Inference
Generating \textit{descriptive} sentences that convey non-trivial, detailed, and salient information about images is an important goal of image captioning. In this paper we propose a novel approach to encourage captionin…
DescriptiveImage CaptioningNatural Language Inference