paper-with-me

홈 › Papers

Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation

2025-05-11 · Xilin Jiang, Junkai Wu, Vishal Choudhari, Nima Mesgarani

Audio large language models (LLMs) are considered experts at recognizing sound objects, yet their performance relative to LLMs in other sensory modalities, such as visual or audio-visual LLMs, and to humans using their ears, eyes, or both remains unexplored. To investigate this, we systematically evaluate audio, visual, and audio-visual LLMs, specifically Qwen2-Audio, Qwen2-VL, and Qwen2.5-Omni, against humans in recognizing sound objects of different classes from audio-only, silent video, or sounded video inputs. We uncover a performance gap between Qwen2-Audio and Qwen2-VL that parallels the sensory discrepancy between human ears and eyes. To reduce this gap, we introduce a cross-modal distillation framework, where an LLM in one modality serves as the teacher and another as the student, with knowledge transfer in sound classes predicted as more challenging to the student by a heuristic model. Distillation in both directions, from Qwen2-VL to Qwen2-Audio and vice versa, leads to notable improvements, particularly in challenging classes. This work highlights the sensory gap in LLMs from a human-aligned perspective and proposes a principled approach to enhancing modality-specific perception in multimodal LLMs.

📄 PDF Abstract BibTeX arXiv:2505.06803

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

Move2Hear: Active Audio-Visual Source Separation

2021-05-15 · ICCV 2021 10 · Sagnik Majumder, Ziad Al-Halah, Kristen Grauman

We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple …

Audio Source SeparationObject

Towards Explainable Quantum AI: Informing the Encoder Selection of Quantum Neural Networks via Visualization

2025-12-16 · Shaolun Ruan, Feng Liang, Rohan Ramakrishna, Chao Ren 외 arxiv

Quantum Neural Networks (QNNs) represent a promising fusion of quantum computing and neural network architectures, offering speed-ups and efficient processing of high-dimensional, entangled data. A crucial component of Q…

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

2025-11-13 · Qilang Ye, Wei Zeng, Meng Liu, Jie Zhang 외 arxiv

Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an ``Audio-Visual Confusion'…

Audio-visual Question AnsweringReinforcement LearningVisual Reasoning

Understanding Audiovisual Deepfake Detection: Techniques, Challenges, Human Factors and Perceptual Insights

2024-11-12 · Ammarah Hashmi, Sahibzada Adil Shahzad, Chia-Wen Lin, Yu Tsao 외

Deep Learning has been successfully applied in diverse fields, and its impact on deepfake detection is no exception. Deepfakes are fake yet realistic synthetic content that can be used deceitfully for political impersona…

DeepFake DetectionFace SwappingMisinformationVideo Forensics

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

2024-11-24 · Luis Vilaca, Yi Yu, Paula Vinan

Audio-visual correlation learning aims to capture and understand natural phenomena between audio and visual data. The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data an…