Hand-Object Interaction and Precise Localization in Transitive Action Recognition
Action recognition in still images has seen major improvement in recent years due to advances in human pose estimation, object recognition and stronger feature representations produced by deep neural networks. However, there are still many cases in which performance remains far from that of humans. A major difficulty arises in distinguishing between transitive actions in which the overall actor pose is similar, and recognition therefore depends on details of the grasp and the object, which may be largely occluded. In this paper we demonstrate how recognition is improved by obtaining precise localization of the action-object and consequently extracting details of the object shape together with the actor-object interaction. To obtain exact localization of the action object and its interaction with the actor, we employ a coarse-to-fine approach which combines semantic segmentation and contextual features, in successive stages. We focus on (but are not limited) to face-related actions, a set of actions that includes several currently challenging categories. We present an average relative improvement of 35% over state-of-the art and validate through experimentation the effectiveness of our approach.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAction Recognition In Still ImagesObjectObject RecognitionPose EstimationSemantic SegmentationTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Hyper-relationship Learning Network for Scene Graph Generation
Generating informative scene graphs from images requires integrating and reasoning from various graph components, i.e., objects and relationships. However, current scene graph generation (SGG) methods, including the unbi…
Graph AttentionGraph GenerationScene Graph GenerationBack to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning
A recurring challenge in preference fine-tuning (PFT) is handling $\textit{intransitive}$ (i.e., cyclic) preferences. Intransitive preferences often stem from either $\textit{(i)}$ inconsistent rankings along a single ob…
Instruction FollowingWeakly supervised training of deep convolutional neural networks for overhead pedestrian localization in depth fields
Overhead depth map measurements capture sufficient amount of information to enable human experts to track pedestrians accurately. However, fully automating this process using image analysis algorithms can be challenging.…
Data AugmentationObject LocalizationOpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggle with generalization, performing well o…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelAlgebraic Semantics of Proto-Transitive Rough Sets
Rough sets over generalized transitive relations like proto-transitive ones had been initiated by the present author in the year 2012. Subsequently, approximation of proto-transitive relations by other relations was inve…