paper-with-me

홈 › Papers

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

2024-06-17 · Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha

Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM) with Advanced Audio Understanding and Complex Reasoning Abilities. We build GAMA by integrating an LLM with multiple types of audio representations, including features from a custom Audio Q-Former, a multi-layer aggregator that aggregates features from multiple layers of an audio encoder. We fine-tune GAMA on a large-scale audio-language dataset, which augments it with audio understanding capabilities. Next, we propose CompA-R (Instruction-Tuning for Complex Audio Reasoning), a synthetically generated instruction-tuning (IT) dataset with instructions that require the model to perform complex reasoning on the input audio. We instruction-tune GAMA with CompA-R to endow it with complex reasoning abilities, where we further add a soft prompt as input with high-level semantic evidence by leveraging event tags of the input audio. Finally, we also propose CompA-R-test, a human-labeled evaluation dataset for evaluating the capabilities of LALMs on open-ended audio question-answering that requires complex reasoning. Through automated and expert human evaluations, we show that GAMA outperforms all other LALMs in literature on diverse audio understanding tasks by margins of 1%-84%. Further, GAMA IT-ed on CompA-R proves to be superior in its complex reasoning and instruction following capabilities.

📄 PDF Abstract BibTeX arXiv:2406.11768

Code (2)

Sreyan88/GAMA jax
sreyan88/reclap pytorch

Tasks

Audio Question AnsweringInstruction FollowingLanguage ModelingLanguage ModellingQuestion Answering

Similar Papers 제목 키워드 기반

MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

2024-10-24 · S Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth 외

The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal au…

Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

2025-03-06 · Sreyan Ghosh, Zhifeng Kong, Sonal Kumar, S Sakshi 외

Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments. In this paper, we introduce Audio Flamingo 2 (AF2), an Audio-Languag…

Audio captioningLanguage ModelingLanguage ModellingQuestion Answering+1

Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference

2024-09-18 · Edresson Casanova, Ryan Langman, Paarth Neekhara, Shehzeen Hussain 외

Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, aud…

Audio CompressionLanguage ModelingLanguage ModellingQuantization+2

SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation

2025-05-06 · Zhaoxi Mu, Xinyu Yang, Gang Wang

While contemporary speech separation technologies adeptly process lengthy mixed audio waveforms, they are frequently challenged by the intricacies of real-world environments, including noisy and reverberant settings, whi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge Distillationspeech-recognition+2

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

2025-06-10 · Ailin Huang, Bingxin Li, Bruce Wang, Boyong Wu 외

Large Audio-Language Models (LALMs) have significantly advanced intelligent human-computer interaction, yet their reliance on text-based outputs limits their ability to generate natural speech responses directly, hinderi…

Language ModelingLanguage ModellingSpeech Synthesis