paper-with-me

Papers

Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

2024-09-27 · Yiming Chen, Xianghu Yue, Xiaoxue Gao, Chen Zhang, Luis Fernando D'Haro, Robby T. Tan, Haizhou Li

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. To bridge this gap, we propose the first multi-audio evaluation (MAE) benchmark that consists of 20 datasets from 11 multi-audio tasks encompassing both speech and sound scenarios. Comprehensive experiments on MAE demonstrate that the existing ALLMs, while being powerful in comprehending primary audio elements in individual audio inputs, struggling to handle multi-audio scenarios. To this end, we propose a novel multi-audio-LLM (MALLM) to capture audio context among multiple similar audios using discriminative learning on our proposed synthetic data. The results demonstrate that the proposed MALLM outperforms all baselines and achieves high data efficiency using synthetic data without requiring human annotations. The proposed MALLM opens the door for ALLMs towards multi-audio processing era and brings us closer to replicating human auditory capabilities in machines.

📄 PDF Abstract BibTeX arXiv:2409.18680

Code (1)

MatthewCYM/MALLM 공식 구현

Methods 이 논문이 사용한 방법론

MAE 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Advancing Audio-Visual Navigation Through Multi-Agent Collaboration in 3D Environments

2025-09-21 · Hailong Zhang, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

Intelligent agents often require collaborative strategies to achieve complex tasks beyond individual capabilities in real-world scenarios. While existing audio-visual navigation (AVN) research mainly focuses on single-ag…

Spatial ReasoningVisual Navigation

TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics

2025-09-30 · Yi-Cheng Lin, Yu-Hua Chen, Jia-Kai Dong, Yueh-Hsuan Huang 외 arxiv

Large audio-language models are advancing rapidly, yet most evaluations emphasize speech or globally sourced sounds, overlooking culturally distinctive cues. This gap raises a critical question: can current models genera…

Question Generation

Audio-Visual Speech Recognition based on Regulated Transformer and Spatio-Temporal Fusion Strategy for Driver Assistive Systems

2024-05-09 · Expert Systems with Applications 2024 5 · Dmitry Ryumin, Alexandr Axyonov, Elena Ryumina, Denis Ivanko 외

This article presents a research methodology for audio-visual speech recognition (AVSR) in driver assistive systems. These systems necessitate ongoing interaction with drivers while driving through voice control for safe…

Audio-Visual Speech RecognitionLipreadingLip Readingspeech-recognition+2

When Misinformation Speaks and Converses: Rethinking Fact-Checking in Audio Platforms

2026-04-18 · Chaewan Chun, Delvin Ce Zhang, Dongwon Lee arxiv

Audio platforms have evolved beyond entertainment. They have become central to public discourse, from podcasts and radio to WhatsApp voice notes and live streams. With millions of shows and hundreds of millions of listen…

The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents

2026-02-15 · Ziyang Ma, Ruiyang Xu, Yinghao Ma, Chao-Han Huck Yang 외 arxiv

Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Challenge at Interspeech 2026, the first sh…

Reinforcement Learning