paper-with-me

Papers

DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment

2024-06-27 · Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, He Huang, Boris Ginsburg, Yu-Chiang Frank Wang, Hung-Yi Lee

Recent speech language models (SLMs) typically incorporate pre-trained speech models to extend the capabilities from large language models (LLMs). In this paper, we propose a Descriptive Speech-Text Alignment approach that leverages speech captioning to bridge the gap between speech and text modalities, enabling SLMs to interpret and generate comprehensive natural language descriptions, thereby facilitating the capability to understand both linguistic and non-linguistic features in speech. Enhanced with the proposed approach, our model demonstrates superior performance on the Dynamic-SUPERB benchmark, particularly in generalizing to unseen tasks. Moreover, we discover that the aligned model exhibits a zero-shot instruction-following capability without explicit speech instruction tuning. These findings highlight the potential to reshape instruction-following SLMs by incorporating rich, descriptive speech captions.

📄 PDF Abstract BibTeX arXiv:2406.18871

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveInstruction Following

Similar Papers 제목 키워드 기반

SpeakGer: A meta-data enriched speech corpus of German state and federal parliaments

2024-10-23 · Kai-Robin Lange, Carsten Jentsch

The application of natural language processing on political texts as well as speeches has become increasingly relevant in political sciences due to the ability to analyze large text corpora which cannot be read by a sing…

DescriptiveSentiment Analysis

DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment

2025-07-03 · Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, Chao-Han Huck Yang 외

We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following, without requiring task-specific audio instruction-tuning. Recent LALMs t…

cross-modal alignmentInstruction FollowingLanguage ModelingLanguage Modelling+1

Multilingual and Explainable Text Detoxification with Parallel Corpora

2024-12-16 · Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele 외

Even with various regulations in place across countries and social media platforms (Government of India, 2021; European Parliament and Council of the European Union, 2022, digital abusive speech remains a significant iss…

DescriptiveStyle TransferText Style Transfer

Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles

2025-10-03 · Rongchen Guo, Vincent Francoeur, Isar Nejadgholi, Sylvain Gagnon 외 arxiv

Speech Emotion Recognition (SER) is essential for improving human-computer interaction, yet its accuracy remains constrained by the complexity of emotional nuances in speech. In this study, we distinguish between descrip…

Speech Emotion Recognition

Enhancing Descriptive Image Captioning with Natural Language Inference

2021-08-01 · ACL 2021 5 · Zhan Shi, Hui Liu, Xiaodan Zhu

Generating \textit{descriptive} sentences that convey non-trivial, detailed, and salient information about images is an important goal of image captioning. In this paper we propose a novel approach to encourage captionin…

DescriptiveImage CaptioningNatural Language Inference