Interaction-aware Joint Attention Estimation Using People Attributes
This paper proposes joint attention estimation in a single image. Different from related work in which only the gaze-related attributes of people are independently employed, (I) their locations and actions are also employed as contextual cues for weighting their attributes, and (ii) interactions among all of these attributes are explicitly modeled in our method. For the interaction modeling, we propose a novel Transformer-based attention network to encode joint attention as low-dimensional features. We introduce a specialized MLP head with positional embedding to the Transformer so that it predicts pixelwise confidence of joint attention for generating the confidence heatmap. This pixelwise prediction improves the heatmap accuracy by avoiding the ill-posed problem in which the high-dimensional heatmap is predicted from the low-dimensional features. The estimated joint attention is further improved by being integrated with general image-based attention estimation. Our method outperforms SOTA methods quantitatively in comparative experiments. Code: https://anonymous.4open.science/r/anonymized_codes-ECA4.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Joint-Relation Transformer for Multi-Person Motion Prediction
Multi-person motion prediction is a challenging problem due to the dependency of motion on both individual past movements and interactions with other people. Transformer-based methods have shown promising results on this…
motion predictionPredictionRelationPandaNet: Anchor-Based Single-Shot Multi-Person 3D Pose Estimation
Recently, several deep learning models have been proposed for 3D human pose estimation. Nevertheless, most of these approaches only focus on the single-person case or estimate 3D pose of a few people at high resolution. …
3D Human Pose Estimation3D Pose EstimationAutonomous DrivingPose EstimationPandaNet : Anchor-Based Single-Shot Multi-Person 3D Pose Estimation
Recently, several deep learning models have been proposed for 3D human pose estimation. Nevertheless, most of these approaches only focus on the single-person case or estimate 3D pose of a few people at high resolution. …
3D Human Pose Estimation3D Pose EstimationAutonomous DrivingPose EstimationInteraction-Aware Topic Model for Microblog Conversations through Network Embedding and User Attention
Traditional topic models are insufficient for topic extraction in social media. The existing methods only consider text information or simultaneously model the posts and the static characteristics of social media. They i…
Network EmbeddingRepresentation LearningTopic ModelsVariational InferenceInvestigating Role of Big Five Personality Traits in Audio-Visual Rapport Estimation
Automatic rapport estimation in social interactions is a central component of affective computing. Recent reports have shown that the estimation performance of rapport in initial interactions can be improved by using the…