paper-with-me

Papers

Audio-Visual Intelligence in Large Foundation Models

2026-05-05 · You Qin, Kai Liu, Shengqiong Wu, Kai Wang, Shijian Deng, Yapeng Tian, Junbin Xiao, Yazhou Xing, Yinghao Ma, Bobo Li, Roger Zimmermann, Lei Cui, Furu Wei, Jiebo Luo, Hao Fei arxiv

Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of large foundation models, joint modeling of audio and vision has become increasingly crucial, i.e., not only for understanding but also for controllable generation and reasoning across dynamic, temporally grounded signals. Recent advances, such as Meta MovieGen and Google Veo-3, highlight the growing industrial and academic focus on unified audio-vision architectures that learn from massive multimodal data. However, despite rapid progress, the literature remains fragmented, spanning diverse tasks, inconsistent taxonomies, and heterogeneous evaluation practices that impede systematic comparison and knowledge integration. This survey provides the first comprehensive review of AVI through the lens of large foundation models. We establish a unified taxonomy covering the broad landscape of AVI tasks, ranging from understanding (e.g., speech recognition, sound localization) to generation (e.g., audio-driven video synthesis, video-to-audio) and interaction (e.g., dialogue, embodied, or agentic interfaces). We synthesize methodological foundations, including modality tokenization, cross-modal fusion, autoregressive and diffusion-based generation, large-scale pretraining, instruction alignment, and preference optimization. Furthermore, we curate representative datasets, benchmarks, and evaluation metrics, offering a structured comparison across task families and identifying open challenges in synchronization, spatial reasoning, controllability, and safety. By consolidating this rapidly expanding field into a coherent framework, this survey aims to serve as a foundational reference for future research on large-scale AVI.

📄 PDF Abstract BibTeX arXiv:2605.04045

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionSpatial Reasoning

Similar Papers 제목 키워드 기반

V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models

2023-08-18 · Heng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright 외

Building artificial intelligence (AI) systems on top of a set of foundation models (FMs) is becoming a new paradigm in AI research. Their representative and generative abilities learnt from vast amounts of data can be ea…

Audio GenerationVideo-to-Sound Generation

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

2025-03-12 · Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

Integrating audio and visual data for training multimodal foundational models remains challenging. We present Audio-Video Vector Alignment (AVVA), which aligns audiovisual (AV) scene content beyond mere temporal synchron…

AudioCapsContrastive LearningLarge Language ModelRetrieval+1

BAVS: Bootstrapping Audio-Visual Segmentation by Integrating Foundation Knowledge

2023-08-20 · Chen Liu, Peike Li, Hu Zhang, Lincheng Li 외

Given an audio-visual pair, audio-visual segmentation (AVS) aims to locate sounding sources by predicting pixel-wise maps. Previous methods assume that each sound component in an audio signal always has a visual counterp…

Audio ClassificationSegmentation

OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models

2025-04-30 · Shengkai Chen, Yifang Yin, Jinming Cao, Shili Xiang 외

Audio-visual segmentation aims to separate sounding objects from videos by predicting pixel-level masks based on audio signals. Existing methods primarily concentrate on closed-set scenarios and direct audio-visual align…

Pseudo LabelSemantic SegmentationTransfer Learning

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

2026-07-17 · Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze 외 arxiv

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLM…

Visual Reasoning