Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
Human-Object Interaction (HOI) detection aims to localize human-object pairs and comprehend their interactions. Recently, two-stage transformer-based methods have demonstrated competitive performance. However, these methods frequently focus on object appearance features and ignore global contextual information. Besides, vision-language model CLIP which effectively aligns visual and text embeddings has shown great potential in zero-shot HOI detection. Based on the former facts, We introduce a novel HOI detector named ISA-HOI, which extensively leverages knowledge from CLIP, aligning interactive semantics between visual and textual features. We first extract global context of image and local features of object to Improve interaction Features in images (IF). On the other hand, we propose a Verb Semantic Improvement (VSI) module to enhance textual features of verb labels via cross-modal fusion. Ultimately, our method achieves competitive results on the HICO-DET and V-COCO benchmarks with much fewer training epochs, and outperforms the state-of-the-art under zero-shot settings.
Code (0)
등록된 구현이 없습니다.
Tasks
Human-Object Interaction DetectionLanguage ModelingLanguage ModellingObjectMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
RR-Net: Injecting Interactive Semantics in Human-Object Interaction Detection
Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects. Latest end-to-end HOI detectors are short of relation reasoning, which leads to inability to learn HOI-specific inte…
Human-Object Interaction DetectionRelationRe-mine, Learn and Reason: Exploring the Cross-modal Semantic Correlations for Language-guided HOI detection
Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predict HOI triplets. Despite the …
Human-Object Interaction DetectionSentenceTransfer LearningSemantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction
Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking. While Large Language Models (LLMs) can automate narrative generation from spatial layouts, current collage-based …
Multi-Semantic Interactive Learning for Object Detection
Single-branch object detection methods use shared features for localization and classification, yet the shared features are not fit for the two different tasks simultaneously. Multi-branch object detection methods usuall…
ClassificationObjectobject-detectionObject Detection+1Exploring Crosslinguistic Frame Alignment
The FrameNet (FN) project at the International Computer Science Institute in Berkeley (ICSI), which documents the core vocabulary of contemporary English, was the first lexical resource based on Fillmore{'}s theory of Fr…