Object State Change Classification in Egocentric Videos using the Divided Space-Time Attention Mechanism
This report describes our submission called "TarHeels" for the Ego4D: Object State Change Classification Challenge. We use a transformer-based video recognition model and leverage the Divided Space-Time Attention mechanism for classifying object state change in egocentric videos. Our submission achieves the second-best performance in the challenge. Furthermore, we perform an ablation study to show that identifying object state change in egocentric videos requires temporal modeling ability. Lastly, we present several positive and negative examples to visualize our model's predictions. The code is publicly available at: https://github.com/md-mohaiminul/ObjectStateChange
Code (1)
Tasks
ObjectObject State Change ClassificationVideo RecognitionSimilar Papers 제목 키워드 기반
Learning State-Aware Visual Representations from Audible Interactions
We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their da…
Action AnticipationAction RecognitionLong Term Action AnticipationObject State Change Classification+1Jointly Recognizing Object Fluents and Tasks in Egocentric Videos
This paper addresses the problem of jointly recognizing object fluents and tasks in egocentric videos. Fluents are the changeable attributes of objects. Tasks are goal-oriented human activities which interact with objec…
ObjectA Semi-Automated Method for Object Segmentation in Infant's Egocentric Videos to Study Object Perception
Object segmentation in infant's egocentric videos is a fundamental step in studying how children perceive objects in early stages of development. From the computer vision perspective, object segmentation in such videos p…
ObjectOptical Flow EstimationSegmentationSemantic SegmentationFine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications
Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labe…
Activity RecognitionData AugmentationObjectSegmentation+2Compact CNN for Indexing Egocentric Videos
While egocentric video is becoming increasingly popular, browsing it is very difficult. In this paper we present a compact 3D Convolutional Neural Network (CNN) architecture for long-term activity recognition in egocentr…
Activity RecognitionOptical Flow Estimation