paper-with-me

홈 › Papers

DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment

2025-07-03 · Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, Chao-Han Huck Yang, Sung-Feng Huang, Chih-Kai Yang, Chee-En Yu, Chun-Wei Chen, Wei-Chih Chen, Chien-yu Huang, Yi-Cheng Lin, Yu-Xiang Lin, Chi-An Fu, Chun-Yi Kuan, Wenze Ren, Xuanjun Chen, Wei-Ping Huang, En-Pei Hu, Tzu-Quan Lin, Yuan-Kuei Wu, Kuan-Po Huang, Hsiao-Ying Huang, Huang-Cheng Chou, Kai-Wei Chang, Cheng-Han Chiang, Boris Ginsburg, Yu-Chiang Frank Wang, Hung-Yi Lee

We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following, without requiring task-specific audio instruction-tuning. Recent LALMs typically augment Large Language Models (LLMs) with auditory capabilities by training on large-scale, manually curated or LLM-synthesized audio-instruction datasets. However, these approaches have often suffered from the catastrophic forgetting of the LLM's original language abilities. To address this, we revisit the data construction pipeline and propose DeSTA, a self-generated cross-modal alignment strategy in which the backbone LLM generates its own training targets. This approach preserves the LLM's native language proficiency while establishing effective audio-text alignment, thereby enabling zero-shot generalization without task-specific tuning. Using DeSTA, we construct DeSTA-AQA5M, a large-scale, task-agnostic dataset containing 5 million training samples derived from 7,000 hours of audio spanning 50 diverse datasets, including speech, environmental sounds, and music. DeSTA2.5-Audio achieves state-of-the-art or competitive performance across a wide range of audio-language benchmarks, including Dynamic-SUPERB, MMAU, SAKURA, Speech-IFEval, and VoiceBench. Comprehensive comparative studies demonstrate that our self-generated strategy outperforms widely adopted data construction and training strategies in both auditory perception and instruction-following capabilities. Our findings underscore the importance of carefully designed data construction in LALM development and offer practical insights for building robust, general-purpose LALMs.

📄 PDF Abstract BibTeX arXiv:2507.02768

Code (2)

kehanlu/desta2.5-audio 공식 구현 pytorch
kehanlu/DeSTA2 pytorch

Tasks

cross-modal alignmentInstruction FollowingLanguage ModelingLanguage ModellingZero-shot Generalization

Similar Papers 제목 키워드 기반

Causal Tracing of Audio-Text Fusion in Large Audio Language Models

2026-03-14 · Wei-Chih Chen, Chien-yu Huang, Hung-yi Lee arxiv

Despite the strong performance of large audio language models (LALMs) in various tasks, exactly how and where they integrate acoustic features with textual context remains unclear. We adapt causal tracing to investigate …

Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning

2025-10-13 · Kuan-Yi Lee, Tsung-En Lin, Hung-Yi Lee arxiv

Recent advancements in large multimodal models (LMMs) have shown strong capabilities in audio understanding. However, most systems rely solely on end-to-end reasoning, limiting interpretability and accuracy for tasks tha…

ASR Bundestag: A Large-Scale political debate dataset in German

2023-02-12 · Johannes Wirth, René Peinl

We present ASR Bundestag, a dataset for automatic speech recognition in German, consisting of 610 hours of aligned audio-transcript pairs for supervised training as well as 1,038 hours of unlabeled audio snippets for sel…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognition+1

Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding

2024-02-05 · Yasar Abbas Ur Rehman, Kin Wai Lau, Yuyang Xie, Lan Ma 외

The integration of Federated Learning (FL) and Self-supervised Learning (SSL) offers a unique and synergetic combination to exploit the audio data for general-purpose audio understanding, without compromising user data p…

Federated LearningRetrievalSelf-Supervised Learning

General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline

2018-07-26 · Eduardo Fonseca, Manoj Plakal, Frederic Font, Daniel P. W. Ellis 외

This paper describes Task 2 of the DCASE 2018 Challenge, titled "General-purpose audio tagging of Freesound content with AudioSet labels". This task was hosted on the Kaggle platform as "Freesound General-Purpose Audio T…

Audio TaggingTask 2