paper-with-me

Papers

Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model

2024-04-19 · Jihao Dong, Renjie Pan, Hua Yang

Human-Object Interaction (HOI) detection aims to localize human-object pairs and comprehend their interactions. Recently, two-stage transformer-based methods have demonstrated competitive performance. However, these methods frequently focus on object appearance features and ignore global contextual information. Besides, vision-language model CLIP which effectively aligns visual and text embeddings has shown great potential in zero-shot HOI detection. Based on the former facts, We introduce a novel HOI detector named ISA-HOI, which extensively leverages knowledge from CLIP, aligning interactive semantics between visual and textual features. We first extract global context of image and local features of object to Improve interaction Features in images (IF). On the other hand, we propose a Verb Semantic Improvement (VSI) module to enhance textual features of verb labels via cross-modal fusion. Ultimately, our method achieves competitive results on the HICO-DET and V-COCO benchmarks with much fewer training epochs, and outperforms the state-of-the-art under zero-shot settings.

📄 PDF Abstract BibTeX arXiv:2404.12678

Code (0)

등록된 구현이 없습니다.

Tasks

Human-Object Interaction DetectionLanguage ModelingLanguage ModellingObject

Methods 이 논문이 사용한 방법론

Focus 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

RR-Net: Injecting Interactive Semantics in Human-Object Interaction Detection

2021-04-30 · Dongming Yang, Yuexian Zou, Can Zhang, Meng Cao 외

Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects. Latest end-to-end HOI detectors are short of relation reasoning, which leads to inability to learn HOI-specific inte…

Human-Object Interaction DetectionRelation

Re-mine, Learn and Reason: Exploring the Cross-modal Semantic Correlations for Language-guided HOI detection

2023-07-25 · ICCV 2023 1 · Yichao Cao, Qingfei Tang, Feng Yang, Xiu Su 외

Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predict HOI triplets. Despite the …

Human-Object Interaction DetectionSentenceTransfer Learning

Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction

2026-04-21 · Xuxin Tang, Ibrahim Tahmid, Eric Krokos, Kirsten Whitley 외 arxiv

Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking. While Large Language Models (LLMs) can automate narrative generation from spatial layouts, current collage-based …

Multi-Semantic Interactive Learning for Object Detection

2023-03-18 · Shuxin Wang, Zhichao Zheng, Yanhui Gu, Junsheng Zhou 외

Single-branch object detection methods use shared features for localization and classification, yet the shared features are not fit for the two different tasks simultaneously. Multi-branch object detection methods usuall…

ClassificationObjectobject-detectionObject Detection+1

Exploring Crosslinguistic Frame Alignment

2020-05-01 · LREC 2020 5 · Collin F. Baker, Arthur Lorenzi

The FrameNet (FN) project at the International Computer Science Institute in Berkeley (ICSI), which documents the core vocabulary of contemporary English, was the first lexical resource based on Fillmore{'}s theory of Fr…