Learning person-object interactions for action recognition in still images
We investigate a discriminatively trained model of person-object interactions for recognizing common human actions in still images. We build on the locally order-less spatial pyramid bag-of-features model, which was shown to perform extremely well on a range of object, scene and human action recognition tasks. We introduce three principal contributions. First, we replace the standard quantized local HOG/SIFT features with stronger discriminatively trained body part and object detectors. Second, we introduce new person-object interaction features based on spatial co-occurrences of individual body parts and objects. Third, we address the combinatorial problem of a large number of possible interaction pairs and propose a discriminative selection procedure using a linear support vector machine (SVM) with a sparsity inducing regularizer. Learning of action-specific body part and object interactions bypasses the difficult problem of estimating the complete human body pose configuration. Benefits of the proposed model are shown on human action recognition in consumer photographs, outperforming the strong bag-of-features baseline.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAction Recognition In Still ImagesObjectTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Face-space Action Recognition by Face-Object Interactions
Action recognition in still images has seen major improvement in recent years due to advances in human pose estimation, object recognition and stronger feature representations. However, there are still many cases in whic…
Action RecognitionAction Recognition In Still ImagesObjectObject Recognition+2Multi-Granularity Reasoning for Social Relation Recognition from Images
Discovering social relations in images can make machines better interpret the behavior of human beings. However, automatically recognizing social relations in images is a challenging task due to the significant gap betwe…
RelationVisual Social Relationship RecognitionGeometric Visual Fusion Graph Neural Networks for Multi-Person Human-Object Interaction Recognition in Videos
Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual f…
Graph LearningGraph Neural NetworkHuman-Object Interaction DetectionFirst-Person Activity Recognition: What Are They Doing to Me?
This paper discusses the problem of recognizing interaction-level human activities from a first-person viewpoint. The goal is to enable an observer (e.g., a robot or a wearable camera) to understand 'what activity others…
Activity RecognitionGeneral ClassificationTHORN: Temporal Human-Object Relation Network for Action Recognition
Most action recognition models treat human activities as unitary events. However, human activities often follow a certain hierarchy. In fact, many human activities are compositional. Also, these actions are mostly human-…
Action RecognitionHuman-Object Interaction DetectionObjectRelation+1