paper-with-me

Papers

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

2024-11-29 · CVPR 2025 1 · Tiantian Geng, Jinrui Zhang, Qingni Wang, Teng Wang, Jinming Duan, Feng Zheng

Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-world videos encompass omni-modal information (vision, audio, and speech) with a series of events forming a cohesive storyline. The lack of multi-modal video data with fine-grained event annotations and the high cost of manual labeling are major obstacles to comprehensive omni-modality video perception. To address this gap, we propose an automatic pipeline consisting of high-quality multi-modal video filtering, semantically coherent omni-modal event boundary detection, and cross-modal correlation-aware event captioning. In this way, we present LongVALE, the first-ever Vision-Audio-Language Event understanding benchmark comprising 105K omni-modal events with precise temporal boundaries and detailed relation-aware captions within 8.4K high-quality long videos. Further, we build a baseline that leverages LongVALE to enable video large language models (LLMs) for omni-modality fine-grained temporal video understanding for the first time. Extensive experiments demonstrate the effectiveness and great potential of LongVALE in advancing comprehensive multi-modal video understanding.

📄 PDF Abstract BibTeX arXiv:2411.19772

Code (1)

ttgeng233/LongVALE 공식 구현 pytorch

Tasks

Boundary DetectionDense Video CaptioningVideo Understanding

Similar Papers 제목 키워드 기반

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

2026-07-15 · Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim 외 arxiv

Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly e…

Instruction Following

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

2026-06-12 · Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida, Kim Sung-Bin 외 arxiv

Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track sou…

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

2026-07-05 · Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang 외 arxiv

Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At t…

Reinforcement LearningSound Event Detection

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

2025-12-14 · Jianghan Chao, Jianzhang Gao, Wenhui Tan, Yuchong Sun 외 arxiv

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including…

Visual Reasoning

TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining

2025-05-12 · Paul Primus, Florian Schmid, Gerhard Widmer

Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text-conditioned audio generation. Existing …

Audio captioningAudio GenerationSentencezero-shot-classification+1