paper-with-me

홈 › Papers

VGGSounder: Audio-Visual Evaluations for Foundation Models

2025-08-11 · Daniil Zverev, Thaddäus Wiedemer, Ameya Prabhu, Matthias Bethge, Wieland Brendel, A. Sophia Koepke arxiv

The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.

📄 PDF Abstract BibTeX arXiv:2508.08237

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

2025-03-12 · Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

Integrating audio and visual data for training multimodal foundational models remains challenging. We present Audio-Video Vector Alignment (AVVA), which aligns audiovisual (AV) scene content beyond mere temporal synchron…

AudioCapsContrastive LearningLarge Language ModelRetrieval+1

Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio

2026-06-09 · Avi Gupta, Nilotpal Sinha, Vishnu Raj, Sambuddha Saha 외 arxiv

Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have spurred significant interest across various modalities, the audio-v…

class-incremental learning

LTX-2: Efficient Joint Audio-Visual Foundation Model

2026-01-06 · Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman 외 arxiv

Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source found…

Audio GenerationVideo Generation

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

2025-12-01 · Le Thien Phuc Nguyen, Zhuoran Yu, Samuel Low Yu Hang, Subin An 외 arxiv

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually so…

Text-to-Audio Generation Synchronized with Videos

2024-03-08 · Shentong Mo, Jing Shi, Yapeng Tian

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models…

AudioCapsAudio GenerationContrastive Learning