paper-with-me

홈 › Papers

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

2026-06-01 · Yaoting Wang, Ziyi Zhang, Wenming Tu, Shaoxuan Xu, Wenjie Du, Cheng Liang, Weijun Wang, Yuanchao Li, Guangyao Li, Hao Fei, Yuanchun Li, Henghui Ding, Yunxin Liu arxiv

Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a cognitively inspired benchmark that evaluates Omni-MLLMs across three stages, perception, understanding, and reasoning, through cross-modal tasks requiring joint audio-visual interpretation. This design enables fine-grained diagnosis of model capabilities and failure modes. To further assess robustness beyond familiar domains, we propose AVI-Bench-PriSe, an extension that probes models' primitive audio-visual sensation using unfamiliar, low-semantic stimuli, testing generalization beyond common training distributions. Extensive experiments on both open-source and closed-source models reveal substantial limitations in current Omni-MLLMs. Based on these findings, we present a four-level AVI taxonomy. Overall, AVI-Bench provides a principled evaluation framework to guide the development of more robust and generalizable AVI. Project website: https://fudancvl.github.io/AVI-Bench/

📄 PDF Abstract BibTeX arXiv:2606.07643

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAVERIX: Multimodal Audio-Visual Evaluation Reasoning IndeX

2025-03-27 · Liuyue Xie, George Z. Wei, Avik Kuthiala, Ce Zheng 외

Frontier models have either been language-only or have primarily focused on vision and language modalities. Although recent advancements in models with vision and audio understanding capabilities have shown substantial p…

Decision Making

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

2026-07-17 · Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze 외 arxiv

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLM…

Visual Reasoning

AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark

2018-07-09 · Sören Becker, Johanna Vielhaben, Marcel Ackermann, Klaus-Robert Müller 외

Explainable Artificial Intelligence (XAI) is targeted at understanding how models perform feature selection and derive their classification decisions. This paper explores post-hoc explanations for deep neural networks in…

Audio ClassificationDecision MakingExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)+2

EgoAVU: Egocentric Audio-Visual Understanding

2026-02-05 · Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja 외 arxiv

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labe…

AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models

2023-09-19 · Yuan Tseng, Layne Berry, Yi-Ting Chen, I-Hsiang Chiu 외

Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information. However, current models often focus on a limited set of tasks, and…

audio-visual learningRepresentation Learning