The "something something" video database for learning and evaluating visual common sense
Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual knowledge with natural language, like humans do, is their lack of common sense knowledge about the physical world. Videos, unlike still images, contain a wealth of detailed information about the physical world. However, most labelled video datasets represent high-level concepts rather than detailed physical aspects about actions and scenes. In this work, we describe our ongoing collection of the "something-something" database of video prediction tasks whose solutions require a common sense understanding of the depicted situation. The database currently contains more than 100,000 videos across 174 classes, which are defined as caption-templates. We also describe the challenges in crowd-sourcing this data at scale.
Code (5)
Tasks
Action RecognitionCommon Sense ReasoningGeneral ClassificationVideo PredictionSimilar Papers 제목 키워드 기반
Retro-Actions: Learning 'Close' by Time-Reversing 'Open' Videos
We investigate video transforms that result in class-homogeneous label-transforms. These are video transforms that consistently maintain or modify the labels of all videos in each class. We propose a general approach to …
Data AugmentationVideo RecognitionZero-Shot LearningTemporal Relational Reasoning in Videos
Temporal relational reasoning, the ability to link meaningful transformations of objects or entities over time, is a fundamental property of intelligent species. In this paper, we introduce an effective and interpretable…
Action ClassificationAction RecognitionAction Recognition In VideosActivity Recognition+4SVT: Supertoken Video Transformer for Efficient Video Understanding
Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video content throughout the network without specia…
Video UnderstandingOn the effectiveness of task granularity for transfer learning
We describe a DNN for video classification and captioning, trained end-to-end, with shared features, to solve tasks at different levels of granularity, exploring the link between granularity in a source task and the qual…
ClassificationDiversityGeneral ClassificationTransfer Learning+1Co-training Transformer with Videos and Images Improves Action Recognition
In learning action recognition, models are typically pre-trained on object recognition with images, such as ImageNet, and later fine-tuned on target action recognition with videos. This approach has achieved good empiric…
Action ClassificationAction RecognitionAction Recognition In VideosObject Recognition+1