paper-with-me

홈 › Papers

AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning

2026-02-11 · Liyang Chen, Hongkai Chen, Yujun Cai, Sifan Li, Qingwen Ye, Yiwei Wang arxiv

Large Audio Language Models (LALMs) have demonstrated strong capabilities in audio understanding and reasoning. However, their performance on fine grained auditory perception remains unreliable, and existing approaches largely rely on data intensive training to internalize perceptual abilities. We propose AudioRouter, a reinforcement learning framework that enables LALMs to improve audio understanding by learning when and how to use external audio tools. Rather than tightly coupling tool usage with audio reasoning, AudioRouter formulates tool use as an explicit decision making problem and optimizes a lightweight routing policy while keeping the underlying reasoning model frozen. Experimental results show that AudioRouter achieves substantial improvements on standard audio understanding benchmarks while requiring up to 600x less training data to learn tool usage compared with conventional training paradigms. These findings suggest that learning effective tool usage offers a data efficient and scalable alternative to internalizing perceptual abilities in LALMs.

📄 PDF Abstract BibTeX arXiv:2602.10439

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDecision Making

Similar Papers 제목 키워드 기반

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

2025-10-13 · Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh 외 arxiv

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced mu…

Multimodal ReasoningAudio Generation

Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

2025-03-06 · Sreyan Ghosh, Zhifeng Kong, Sonal Kumar, S Sakshi 외

Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments. In this paper, we introduce Audio Flamingo 2 (AF2), an Audio-Languag…

Audio captioningLanguage ModelingLanguage ModellingQuestion Answering+1

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

2024-06-17 · Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru 외

Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Langu…

Audio Question AnsweringInstruction FollowingLanguage ModelingLanguage Modelling+1

Spatial Audio Motion Understanding and Reasoning

2025-09-18 · Arvind Krishna Sridhar, Yinyi Guo, Erik Visser arxiv

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving…

Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs

2025-06-07 · Wenyu Zhang, Yingxu He, Geyu Lin, Zhuohan Liu 외

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches ofte…

Emotion Recognitionspeech-recognitionSpeech Recognition