paper-with-me

Papers

CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges

2025-09-26 · Hui Li, Changhao Jiang, Hongyu Wang, Ming Zhang, Jiajun Sun, Zhixiong Yang, Yifei Cao, Shihan Dou, Xiaoran Fan, Baoyu Fan, Tao Ji, Tao Gui, Qi Zhang, Xuanjing Huang arxiv

The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene settings and English audio data and do not fully capture scenarios where multiple speakers, unfolding events, and heterogeneous audio sources interact. To address these challenges, we introduce CMDAR, a Chinese benchmark for evaluating models on complex, multi-scene, and dynamically evolving audio reasoning tasks. CMDAR comprises 3,000 carefully curated question-answer pairs linked to diverse audio clips, covering five categories of complex reasoning and spanning three question types. We benchmark 26 state-of-the-art audio language models on CMDAR and observe that they exhibit limitations in complex reasoning tasks. In CMDAR-main, Qwen2.5-Omni achieves 76.67% accuracy, whereas GPT-4o Audio reaches 68.47%. However, GPT-4o Audio substantially outperforms Qwen2.5-Omni on the more challenging multiple-choice with multiple audios and open-ended tasks. And we provide detail analysis corresponding suggestions for the future development of large audio language models.

📄 PDF Abstract BibTeX arXiv:2509.22461

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AudioScenic: Audio-Driven Video Scene Editing

2024-04-25 · Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 외

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image edi…

Characterizing dynamically varying acoustic scenes from egocentric audio recordings in workplace setting

2019-11-10 · Arindam Jati, Amrutha Nadarajan, Karel Mundnich, Shrikanth Narayanan

Devices capable of detecting and categorizing acoustic scenes have numerous applications such as providing context-aware user experiences. In this paper, we address the task of characterizing acoustic scenes in a workpla…

Acoustic Scene ClassificationGeneral ClassificationScene Classification

Be Everywhere - Hear Everything (BEE): Audio Scene Reconstruction by Sparse Audio-Visual Samples

2023-01-01 · ICCV 2023 1 · Mingfei Chen, Kun Su, Eli Shlizerman

Fully immersive and interactive audio-visual scenes are dynamic such that the listeners and the sound emitters move and interact with each other. Reconstruction of an immersive sound experience, as it happens in the …

OCR-Enhanced Multimodal ASR Can Read While Listening

2026-01-26 · Junli Chen, Changli Tang, Yixuan Li, Guangzhi Sun 외 arxiv

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve s…

Audio-Visual Speech RecognitionKnowledge Distillation

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

2025-08-02 · Jui-Ming Yao, Bing-Cheng Xie, Sheng-Wei Peng, Hao-Yuan Chen 외 arxiv

Multimodal Large Language Models (MLLMs) process visual, acoustic, and textual inputs, addressing the limitations of single-modality LLMs. However, existing benchmarks often overlook tri-modal evaluation in Traditional C…

Question Answering