Knowledge Guided Bidirectional Attention Network for Human-Object Interaction Detection
Human Object Interaction (HOI) detection is a challenging task that requires to distinguish the interaction between a human-object pair. Attention based relation parsing is a popular and effective strategy utilized in HOI. However, current methods execute relation parsing in a "bottom-up" manner. We argue that the independent use of the bottom-up parsing strategy in HOI is counter-intuitive and could lead to the diffusion of attention. Therefore, we introduce a novel knowledge-guided top-down attention into HOI, and propose to model the relation parsing as a "look and search" process: execute scene-context modeling (i.e. look), and then, given the knowledge of the target pair, search visual clues for the discrimination of the interaction between the pair. We implement the process via unifying the bottom-up and top-down attention in a single encoder-decoder based model. The experimental results show that our model achieves competitive performance on the V-COCO and HICO-DET datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderHuman-Object Interaction DetectionRelationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Human Attention-Guided Explainable Artificial Intelligence for Computer Vision Models
We examined whether embedding human attention knowledge into saliency-based explainable AI (XAI) methods for computer vision models could enhance their plausibility and faithfulness. We first developed new gradient-based…
ClassificationExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)image-classification+4Attentive Fashion Grammar Network for Fashion Landmark Detection and Clothing Category Classification
This paper proposes a knowledge-guided fashion network to solve the problem of visual fashion analysis, e.g., fashion landmark localization and clothing category classification. The suggested fashion model is leveraged w…
General ClassificationCross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation
Given a natural language expression and a remote sensing image, the goal of referring remote sensing image segmentation (RRSIS) is to generate a pixel-level mask of the target object identified by the referring expressio…
BenchmarkingImage SegmentationReferring ExpressionSemantic SegmentationGCAN: Generative Counterfactual Attention-guided Network for Explainable Cognitive Decline Diagnostics based on fMRI Functional Connectivity
Diagnosis of mild cognitive impairment (MCI) and subjective cognitive decline (SCD) from fMRI functional connectivity (FC) has gained popularity, but most FC-based diagnostic models are black boxes lacking casual reasoni…
counterfactualCounterfactual ReasoningDiagnosticFunctional Connectivity+1MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities
While originally designed for unidirectional generative modeling, decoder-only large language models (LLMs) are increasingly being adapted for bidirectional modeling. However, unidirectional and bidirectional models are …
DecoderLanguage ModelingLanguage ModellingRepresentation Learning+2