paper-with-me

홈 › Papers

FitPro: A Zero-Shot Framework for Interactive Text-based Pedestrian Retrieval in Open World

2025-09-20 · Zengli Luo, Canlong Zhang, Xiaochun Lu, Zhixin Li arxiv

Text-based Pedestrian Retrieval (TPR) deals with retrieving specific target pedestrians in visual scenes according to natural language descriptions. Although existing methods have achieved progress under constrained settings, interactive retrieval in the open-world scenario still suffers from limited model generalization and insufficient semantic understanding. To address these challenges, we propose FitPro, an open-world interactive zero-shot TPR framework with enhanced semantic comprehension and cross-scene adaptability. FitPro has three innovative components: Feature Contrastive Decoding (FCD), Incremental Semantic Mining (ISM), and Query-aware Hierarchical Retrieval (QHR). The FCD integrates prompt-guided contrastive decoding to generate high-quality structured pedestrian descriptions from denoised images, effectively alleviating semantic drift in zero-shot scenarios. The ISM constructs holistic pedestrian representations from multi-view observations to achieve global semantic modeling in multi-turn interactions, thereby improving robustness against viewpoint shifts and fine-grained variations in descriptions. The QHR dynamically optimizes the retrieval pipeline according to query types, enabling efficient adaptation to multi-modal and multi-view inputs. Extensive experiments on five public datasets and two evaluation protocols demonstrate that FitPro significantly overcomes the generalization limitations and semantic modeling constraints of existing methods in interactive retrieval, paving the way for practical deployment.

📄 PDF Abstract BibTeX arXiv:2509.16674

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Temporally-Extended Prompts Optimization for SAM in Interactive Medical Image Segmentation

2023-06-15 · Chuyun Shen, Wenhao Li, Ya zhang, Xiangfeng Wang

The Segmentation Anything Model (SAM) has recently emerged as a foundation model for addressing image segmentation. Owing to the intrinsic complexity of medical images and the high annotation cost, the medical image segm…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection

2024-08-05 · Ting Lei, Shaofeng Yin, Yuxin Peng, Yang Liu

Zero-shot Human-Object Interaction (HOI) detection has emerged as a frontier topic due to its capability to detect HOIs beyond a predefined set of categories. This task entails not only identifying the interactiveness of…

Human-Object Interaction DetectionPrompt Learning

Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model

2024-04-19 · Jihao Dong, Renjie Pan, Hua Yang

Human-Object Interaction (HOI) detection aims to localize human-object pairs and comprehend their interactions. Recently, two-stage transformer-based methods have demonstrated competitive performance. However, these meth…

Human-Object Interaction DetectionLanguage ModelingLanguage ModellingObject

End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge Distillation

2022-04-01 · Mingrui Wu, Jiaxin Gu, Yunhang Shen, Mingbao Lin 외

Most existing Human-Object Interaction~(HOI) Detection methods rely heavily on full annotations with predefined HOI categories, which is limited in diversity and costly to scale further. We aim at advancing zero-shot HOI…

Human-Object Interaction DetectionKnowledge DistillationObjectobject-detection+2

Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language Transfer

2024-03-18 · IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024 3 · Sandipan Sarma, Pradnesh Kalkar, Arijit Sur

Human-Object Interaction (HOI) detection is a crucial task that involves localizing interactive human-object pairs and identifying the actions being performed. Most existing HOI detectors are supervised in nature and lac…

Human-Object Interaction DetectionLanguage ModelingLanguage ModellingObject+1