paper-with-me

홈 › Papers

Weakly-Supervised Action Detection Guided by Audio Narration

2022-05-12 · Keren Ye, Adriana Kovashka

Videos are more well-organized curated data sources for visual concept learning than images. Unlike the 2-dimensional images which only involve the spatial information, the additional temporal dimension bridges and synchronizes multiple modalities. However, in most video detection benchmarks, these additional modalities are not fully utilized. For example, EPIC Kitchens is the largest dataset in first-person (egocentric) vision, yet it still relies on crowdsourced information to refine the action boundaries to provide instance-level action annotations. We explored how to eliminate the expensive annotations in video detection data which provide refined boundaries. We propose a model to learn from the narration supervision and utilize multimodal features, including RGB, motion flow, and ambient sound. Our model learns to attend to the frames related to the narration label while suppressing the irrelevant frames from being used. Our experiments show that noisy audio narration suffices to learn a good action detection model, thus reducing annotation expenses.

📄 PDF Abstract BibTeX arXiv:2205.05895

Code (0)

등록된 구현이 없습니다.

Tasks

Action Detection

Similar Papers 제목 키워드 기반

Guided learning for weakly-labeled semi-supervised sound event detection

2019-06-06 · Liwei Lin, Xiangdong Wang, Hong Liu, Yueliang Qian

We propose a simple but efficient method termed Guided Learning for weakly-labeled semi-supervised sound event detection (SED). There are two sub-targets implied in weakly-labeled SED: audio tagging and boundary detectio…

Audio TaggingBoundary DetectionEvent DetectionGeneral Classification+1

Weakly-supervised Audio-visual Sound Source Detection and Separation

2021-03-25 · Tanzila Rahman, Leonid Sigal

Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mi…

Audio Source SeparationDenoisingObjectSegmentation+3

Past and Future Motion Guided Network for Audio Visual Event Localization

2022-05-08 · Tingxiu Chen, Jianqin Yin, Jin Tang

In recent years, audio-visual event localization has attracted much attention. It's purpose is to detect the segment containing audio-visual events and recognize the event category from untrimmed videos. Existing methods…

audio-visual event localization

Audio-Guided Attention Network for Weakly Supervised Violence Detection

2022-02-21 · Conference 2022 2 · Yujiang Pu, Xiaoyu Wu

Detecting violence in video is a challenging task due to its complex scenarios and great intra-class variability. Most previous works specialize in the analysis of appearance or motion information, ignoring the co-occurr…

Anomaly Detection In Surveillance Videos

Weakly Supervised Scalable Audio Content Analysis

2016-06-12 · Anurag Kumar, Bhiksha Raj

Audio Event Detection is an important task for content analysis of multimedia data. Most of the current works on detection of audio events is driven through supervised learning approaches. We propose a weakly supervised …

Event DetectionMultiple Instance LearningWeakly-supervised Learning