A flexible model for training action localization with varying levels of supervision
Spatio-temporal action detection in videos is typically addressed in a fully-supervised setup with manual annotation of training videos required at every frame. Since such annotation is extremely tedious and prohibits scalability, there is a clear need to minimize the amount of manual supervision. In this work we propose a unifying framework that can handle and combine varying types of less-demanding weak supervision. Our model is based on discriminative clustering and integrates different types of supervision as constraints on the optimization. We investigate applications of such a model to training setups with alternative supervisory signals ranging from video-level class labels to the full per-frame annotation of action bounding boxes. Experiments on the challenging UCF101-24 and DALY datasets demonstrate competitive performance of our method at a fraction of supervision used by previous methods. The flexibility of our model enables joint learning from data with different levels of annotation. Experimental results demonstrate a significant gain by adding a few fully supervised examples to otherwise weakly labeled videos.
Code (1)
Tasks
Action DetectionAction LocalizationClusteringSimilar Papers 제목 키워드 기반
MEG Source Localization via Deep Learning
We present a deep learning solution to the problem of localization of magnetoencephalography (MEG) brain signals. The proposed deep model architectures are tuned for single and multiple time point MEG data, and can estim…
Deep LearningN2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields
Understanding complex scenes at multiple levels of abstraction remains a formidable challenge in computer vision. To address this, we introduce Nested Neural Feature Fields (N2F2), a novel approach that employs hierarchi…
Scene UnderstandingAdaptive Control Attention Network for Underwater Acoustic Localization and Domain Adaptation
Localizing acoustic sound sources in the ocean is a challenging task due to the complex and dynamic nature of the environment. Factors such as high background noise, irregular underwater geometries, and varying acoustic …
Domain AdaptationAnalysis of Contraction Effort Level in EMG-Based Gesture Recognition Using Hyperdimensional Computing
Varying contraction levels of muscles is a big challenge in electromyography-based gesture recognition. Some use cases require the classifier to be robust against varying force changes, while others demand to distinguish…
General ClassificationGesture RecognitionHand Gesture RecognitionHand-Gesture RecognitionReinforcement Learning Constrained Beam Search for Parameter Optimization of Paper Drying Under Flexible Constraints
Existing approaches to enforcing design constraints in Reinforcement Learning (RL) applications often rely on training-time penalties in the reward function or training/inference-time invalid action masking, but these me…
Combinatorial Optimizationreinforcement-learningReinforcement LearningReinforcement Learning (RL)