Action Detection by Implicit Intentional Motion Clustering
Explicitly using human detection and pose estimation has found limited success in action recognition problems. This may be due to the complexity in the articulated motion human exhibit. Yet, we know that action requires an actor and intention. This paper hence seeks to understand the spatiotemporal properties of intentional movement and how to capture such intentional movement without relying on challenging human detection and tracking. We conduct a quantitative analysis of intentional movement, and our findings motivate a new approach for implicit intentional movement extraction that is based on spatiotemporal trajectory clustering by leveraging the properties of intentional movement. The intentional movement clusters are then used as action proposals for detection. Our results on three action detection benchmarks indicate the relevance of focusing on intentional movement for action detection; our method significantly outperforms the state of the art on the challenging MSR-II multi-action video benchmark.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionAction RecognitionClusteringHuman DetectionPose EstimationTemporal Action LocalizationTrajectory ClusteringSimilar Papers 제목 키워드 기반
Leveraging Self-Supervised Training for Unintentional Action Recognition
Unintentional actions are rare occurrences that are difficult to define precisely and that are highly dependent on the temporal context of the action. In this work, we explore such actions and seek to identify the points…
Action RecognitionEmotion and Intention Guided Multi-Modal Learning for Sticker Response Selection
Stickers are widely used in online communication to convey emotions and implicit intentions. The Sticker Response Selection (SRS) task aims to select the most contextually appropriate sticker based on the dialogue. Howev…
Adding Knowledge to Unsupervised Algorithms for the Recognition of Intent
Computer vision algorithms performance are near or superior to humans in the visual problems including object recognition (especially those of fine-grained categories), segmentation, and 3D object reconstruction from 2D …
3D Object ReconstructionObject RecognitionObject ReconstructionDo Robots Need Body Language? Comparing Communication Modalities for Legible Motion Intent in Human-Shared Spaces
Robots in shared spaces often move in ways that are difficult for people to interpret, placing the burden on humans to adapt. High-DoF robots exhibit motion that people read as expressive, intentionally or not, making it…
Contextually learnt detection of unusual motion-based behaviour in crowded public spaces
In this paper we are interested in analyzing behaviour in crowded public places at the level of holistic motion. Our aim is to learn, without user input, strong scene priors or labelled data, the scope of "normal behavio…
Clustering