paper-with-me

Papers

Benchmarking Audio Visual Segmentation for Long-Untrimmed Videos

2024-01-01 · CVPR 2024 1 · Chen Liu, Peike Patrick Li, Qingtao Yu, Hongwei Sheng, Dadong Wang, Lincheng Li, Xin Yu

Existing audio-visual segmentation datasets typically focus on short-trimmed videos with only one pixel-map annotation for a per-second video clip. In contrast for untrimmed videos the sound duration start- and end-sounding time positions and visual deformation of audible objects vary significantly. Therefore we observed that current AVS models trained on trimmed videos might struggle to segment sounding objects in long videos. To investigate the feasibility of grounding audible objects in videos along both temporal and spatial dimensions we introduce the Long-Untrimmed Audio-Visual Segmentation dataset (LU-AVS) which includes precise frame-level annotations of sounding emission times and provides exhaustive mask annotations for all frames. Considering that pixel-level annotations are difficult to achieve in some complex scenes we also provide the bounding boxes to indicate the sounding regions. Specifically LU-AVS contains 10M mask annotations across 6.6K videos and 11M bounding box annotations across 7K videos. Compared with the existing datasets LU-AVS videos are on average 4 8 times longer with the silent duration being 3 15 times greater. Furthermore we try our best to adapt some baseline models that were originally designed for audio-visual-relevant tasks to examine the challenges of our newly curated LU-AVS. Through comprehensive evaluation we demonstrate the challenges of LU-AVS compared to the ones containing trimmed videos. Therefore LU-AVS provides an ideal yet challenging platform for evaluating audio-visual segmentation and localization on untrimmed long videos. The dataset is publicly available at: https://yenanliu.github.io/LU-AVS/.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

APES: Audiovisual Person Search in Untrimmed Video

2021-06-03 · Juan Leon Alcazar, Long Mai, Federico Perazzi, Joon-Young Lee 외

Humans are arguably one of the most important subjects in video streams, many real-world applications such as video summarization or video editing workflows often require the automatic search and retrieval of a person of…

Person RetrievalPerson SearchRetrievalVideo Editing+1

Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline

2023-03-22 · CVPR 2023 1 · Tiantian Geng, Teng Wang, Jinming Duan, Runmin Cong 외

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual event…

audio-visual event localization

Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding

2024-03-24 · Yunlong Tang, Daiki Shimada, Jing Bi, Mingqian Feng 외

Large language models (LLMs) have demonstrated remarkable capabilities in natural language and multimodal domains. By fine-tuning multimodal LLMs with temporal annotations from well-annotated datasets, e.g., dense video …

Dense Video CaptioningTemporal LocalizationVideo CaptioningVideo Understanding

Listen to Look: Action Recognition by Previewing Audio

2019-12-10 · CVPR 2020 6 · Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, Lorenzo Torresani

In the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical. We propose a framework for efficient action recognition in untrimmed video that uses audio as a preview mechani…

Action Recognition

Joint Visual-Temporal Embedding for Unsupervised Learning of Actions in Untrimmed Sequences

2020-01-29 · Rosaura G. VidalMata, Walter J. Scheirer, Anna Kukleva, David Cox 외

Understanding the structure of complex activities in untrimmed videos is a challenging task in the area of action recognition. One problem here is that this task usually requires a large amount of hand-annotated minute- …

Action RecognitionAction SegmentationTemporal Localization