paper-with-me

홈 › Papers

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

2025-08-10 · Siminfar Samakoush Galougah, Rishie Raj, Sanjoy Chowdhury, Sayan Nag, Ramani Duraiswami arxiv

Current audio-visual (AV) benchmarks focus on final answer accuracy, overlooking the underlying reasoning process. This makes it difficult to distinguish genuine comprehension from correct answers derived through flawed reasoning or hallucinations. To address this, we introduce AURA (Audio-visual Understanding and Reasoning Assessment), a benchmark for evaluating the cross-modal reasoning capabilities of Audio-Visual Large Language Models (AV-LLMs) and Omni-modal Language Models (OLMs). AURA includes questions across six challenging cognitive domains, such as causality, timbre and pitch, tempo and AV synchronization, unanswerability, implicit distractions, and skill profiling, explicitly designed to be unanswerable from a single modality. This forces models to construct a valid logical path grounded in both audio and video, setting AURA apart from AV datasets that allow uni-modal shortcuts. To assess reasoning traces, we propose a novel metric, AuraScore, which addresses the lack of robust tools for evaluating reasoning fidelity. It decomposes reasoning into two aspects: (i) Factual Consistency - whether reasoning is grounded in perceptual evidence, and (ii) Core Inference - the logical validity of each reasoning step. Evaluations of SOTA models on AURA reveal a critical reasoning gap: although models achieve high accuracy (up to 92% on some tasks), their Factual Consistency and Core Inference scores fall below 45%. This discrepancy highlights that models often arrive at correct answers through flawed logic, underscoring the need for our benchmark and paving the way for more robust multimodal evaluation.

📄 PDF Abstract BibTeX arXiv:2508.07470

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

KddRES: A Multi-level Knowledge-driven Dialogue Dataset for Restaurant Towards Customized Dialogue System

2020-11-17 · Hongru Wang, Min Li, Zimo Zhou, Gabriel Pui Cheong Fung 외

Compared with CrossWOZ (Chinese) and MultiWOZ (English) dataset which have coarse-grained information, there is no dataset which handle fine-grained and hierarchical level information properly. In this paper, we publish …

Diversity

Cyclic Learning for Binaural Audio Generation and Localization

2024-01-01 · CVPR 2024 1 · Zhaojian Li, Bin Zhao, Yuan Yuan

Binaural audio is obtained by simulating the biological structure of human ears which plays an important role in artificial immersive spaces. A promising approach is to utilize mono audio and corresponding vision to …

Audio GenerationObjectObject Localization

Fine-grained Image Classification by Exploring Bipartite-Graph Labels

2015-12-08 · CVPR 2016 6 · Feng Zhou, Yuanqing Lin

Given a food image, can a fine-grained object recognition engine tell "which restaurant which dish" the food belongs to? Such ultra-fine grained image recognition is the key for many applications like search by images, b…

ClassificationFine-Grained Image ClassificationFine-Grained Image RecognitionGeneral Classification+4

Temporally Aligned Audio for Video with Autoregression

2024-09-20 · Ilpo Viertola, Vladimir Iashin, Esa Rahtu

We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feat…

Audio GenerationVideo-to-Sound Generation

CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation

2025-01-06 · Yuanhong Chen, Kazuki Shimada, Christian Simon, Yukara Ikemiya 외

Binaural audio generation (BAG) aims to convert monaural audio to stereo audio using visual prompts, requiring a deep understanding of spatial and semantic information. However, current models risk overfitting to room en…

Audio GenerationContrastive Learning