paper-with-me

홈 › Papers

Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration

2025-08-05 · Ting Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng, Yang Liu arxiv

Open Vocabulary Human-Object Interaction (HOI) detection aims to detect interactions between humans and objects while generalizing to novel interaction classes beyond the training set. Current methods often rely on Vision and Language Models (VLMs) but face challenges due to suboptimal image encoders, as image-level pre-training does not align well with the fine-grained region-level interaction detection required for HOI. Additionally, effectively encoding textual descriptions of visual appearances remains difficult, limiting the model's ability to capture detailed HOI relationships. To address these issues, we propose INteraction-aware Prompting with Concept Calibration (INP-CC), an end-to-end open-vocabulary HOI detector that integrates interaction-aware prompts and concept calibration. Specifically, we propose an interaction-aware prompt generator that dynamically generates a compact set of prompts based on the input scene, enabling selective sharing among similar interactions. This approach directs the model's attention to key interaction patterns rather than generic image-level semantics, enhancing HOI detection. Furthermore, we refine HOI concept representations through language model-guided calibration, which helps distinguish diverse HOI concepts by investigating visual similarities across categories. A negative sampling strategy is also employed to improve inter-modal similarity modeling, enabling the model to better differentiate visually similar but semantically distinct actions. Extensive experimental results demonstrate that INP-CC significantly outperforms state-of-the-art models on the SWIG-HOI and HICO-DET datasets. Code is available at https://github.com/ltttpku/INP-CC.

📄 PDF Abstract BibTeX arXiv:2508.03207

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Identifying the Unknown: Prompt-Free Open Vocabulary Anomaly Recognition for Robot-Object Interaction

2026-06-25 · Philipp Allgeuer, Jan-Gerrit Habekost, Stefan Wermter arxiv

Robots operating in real-world environments must in general be able to recognize previously unseen objects. As robotic systems move toward open-world autonomy, there is a growing, yet largely unmet, need for open vocabul…

Anomaly Detection

Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

2024-12-26 · Tao Liu, Rongjie Li, Chongyu Wang, Xuming He

Open-vocabulary Scene Graph Generation (OV-SGG) overcomes the limitations of the closed-set assumption by aligning visual relationship representations with open-vocabulary textual representations. This enables the identi…

Graph GenerationLarge Language ModelRelationScene Graph Generation+1

End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting

2024-09-19 · Yongqi Wang, Shuo Yang, Xinxiao wu, Jiebo Luo

Open-vocabulary video visual relationship detection aims to expand video visual relationship detection beyond annotated categories by detecting unseen relationships between both seen and unseen objects in videos. Existin…

DecoderObjectobject-detectionObject Detection+5

RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection

2024-05-30 · Fangyi Chen, Han Zhang, Zhantao Yang, Hao Chen 외

Open-vocabulary object detection (OVD) requires solid modeling of the region-semantic relationship, which could be learned from massive region-text pairs. However, such data is limited in practice due to significant anno…

Image CaptioningImage InpaintingObjectobject-detection+4

OpenSD: Unified Open-Vocabulary Segmentation and Detection

2023-12-10 · Shuai Li, Minghan Li, Pengfei Wang, Lei Zhang

Recently, a few open-vocabulary methods have been proposed by employing a unified architecture to tackle generic segmentation and detection tasks. However, their performance still lags behind the task-specific models due…

DecoderPrompt LearningSegmentationZero Shot Segmentation