paper-with-me

Papers

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

2026-07-01 · Tianhong Zhou, Mingyang Han, Boyu Li, Yuxuan Jiang, Jiaxin Ye, Dongxiao Wang, Haoxiang Shi, Kunpeng Wang, Jun Song, Cheng Yu, Bo Zheng arxiv

Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models exhibit dimensional bias, typically focusing on either semantic matching or temporal offset detection. Moreover, their data construction remains coupled, preventing independent assessment of temporal and semantic consistency. We propose AV-SyncBench, the first benchmark to fully separate temporal and semantic evaluation for audio-visual synchronization. Built from in-the-wild videos, it spans Voice, Music, and Sound across 10 scenarios and 5 challenge tasks. Data are automatically filtered and manually verified to ensure on-screen sound sources. The benchmark contains 3,269 videos and 38,390 samples, and we evaluate five representative models to quantify feature quality for alignment and downstream tasks. The code and dataset are available at: https://fgt7t6g.github.io/AV-SyncBench.

📄 PDF Abstract BibTeX arXiv:2607.00726

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning

2025-05-26 · Wenrui Li, Penghong Wang, Xingtao Wang, WangMeng Zuo 외

Audio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodologies often struggle with background scen…

Zero-Shot Learning

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

2024-12-15 · Pengcheng Zhao, Jinxing Zhou, Yang Zhao, Dan Guo 외

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital.…

Temporal Bilinear Encoding Network of Audio-Visual Features at Low Sampling Rates

2020-12-18 · Feiyan Hu, Eva Mohedano, Noel O'Connor, Kevin McGuinness

Current deep learning based video classification architectures are typically trained end-to-end on large volumes of data and require extensive computational resources. This paper aims to exploit audio-visual information …

ClassificationGeneral ClassificationVideo Classification

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

2026-05-08 · Hengyi Feng, Hao Liang, Mingrui Chen, Bohan Zeng 외 arxiv

Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks largely fail to evaluate this capabilit…

Multimodal Reasoning

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

2026-05-09 · Shihao Cheng, Jiaxu Zhang, Quanyue Song, Shansong Liu 외 arxiv

Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often …

Video Generation