Event Recognition in Videos by Learning from Heterogeneous Web Sources
In this work, we propose to leverage a large number of loosely labeled web videos (e.g., from YouTube) and web images (e.g., from Google/Bing image search) for visual event recognition in consumer videos without requiring any labeled consumer videos. We formulate this task as a new multi-domain adaptation problem with heterogeneous sources, in which the samples from different source domains can be represented by different types of features with different dimensions (e.g., the SIFT features from web images and space-time (ST) features from web videos) while the target domain samples have all types of features. To effectively cope with the heterogeneous sources where some source domains are more relevant to the target domain, we propose a new method called Multi-domain Adaptation with Heterogeneous Sources (MDA-HS) to learn an optimal target classifier, in which we simultaneously seek the optimal weights for different source domains with different types of features as well as infer the labels of unlabeled target domain data based on multiple types of features. We solve our optimization problem by using the cutting-plane algorithm based on group-based multiple kernel learning. Comprehensive experiments on two datasets demonstrate the effectiveness of MDA-HS for event recognition in consumer videos.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationImage RetrievalSimilar Papers 제목 키워드 기반
Human Action Recognition in Drone Videos using a Few Aerial Training Examples
Drones are enabling new forms of human actions surveillance due to their low cost and fast mobility. However, using deep neural networks for automatic aerial action recognition is difficult due to the need for a large nu…
Action ClassificationAction RecognitionTemporal Action LocalizationHeterogeneous Knowledge Transfer in Video Emotion Recognition, Attribution and Summarization
Emotion is a key element in user-generated videos. However, it is difficult to understand emotions conveyed in such videos due to the complex and unstructured nature of user-generated content and the sparsity of video fr…
Emotion RecognitionTransfer LearningVideo Emotion RecognitionZero-Shot LearningOpenLongTail: Generative Scaling of Long-Tail Driving Data
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilize…
Autonomous DrivingRhyRNN: Rhythmic RNN for Recognizing Events in Long and Complex Videos
Though many successful approaches have been proposed for recognizing events in short and homogeneous videos, doing so with long and complex videos remains a challenge. One particular reason is that events in long and com…
RhythmTemporal Stochastic Softmax for 3D CNNs: An Application in Facial Expression Recognition
Training deep learning models for accurate spatiotemporal recognition of facial expressions in videos requires significant computational resources. For practical reasons, 3D Convolutional Neural Networks (3D CNNs) are us…
Facial Expression RecognitionFacial Expression Recognition (FER)