paper-with-me

홈 › Papers

Exploiting the Textual Potential from Vision-Language Pre-training for Text-based Person Search

2023-03-08 · Guanshuo Wang, Fufu Yu, Junjie Li, Qiong Jia, Shouhong Ding

Text-based Person Search (TPS), is targeted on retrieving pedestrians to match text descriptions instead of query images. Recent Vision-Language Pre-training (VLP) models can bring transferable knowledge to downstream TPS tasks, resulting in more efficient performance gains. However, existing TPS methods improved by VLP only utilize pre-trained visual encoders, neglecting the corresponding textual representation and breaking the significant modality alignment learned from large-scale pre-training. In this paper, we explore the full utilization of textual potential from VLP in TPS tasks. We build on the proposed VLP-TPS baseline model, which is the first TPS model with both pre-trained modalities. We propose the Multi-Integrity Description Constraints (MIDC) to enhance the robustness of the textual modality by incorporating different components of fine-grained corpus during training. Inspired by the prompt approach for zero-shot classification with VLP models, we propose the Dynamic Attribute Prompt (DAP) to provide a unified corpus of fine-grained attributes as language hints for the image modality. Extensive experiments show that our proposed TPS framework achieves state-of-the-art performance, exceeding the previous best method by a margin.

📄 PDF Abstract BibTeX arXiv:2303.04497

Code (0)

등록된 구현이 없습니다.

Tasks

AttributePerson SearchText based Person Searchzero-shot-classificationZero-Shot Learning

Similar Papers 제목 키워드 기반

Exploiting Category Names for Few-Shot Classification with Vision-Language Models

2022-11-29 · Taihong Xiao, ZiRui Wang, Liangliang Cao, Jiahui Yu 외

Vision-language foundation models pretrained on large-scale data provide a powerful tool for many visual understanding tasks. Notably, many vision-language models build two encoders (visual and textual) that can map two …

ClassificationFew-Shot Image Classificationimage-classificationImage Classification

Zero-Shot Temporal Action Localization Through Textual Guidance

2026-05-21 · Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Paolo Rota 외 arxiv

Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at training time. Existing work uses Vision and Language Models (VLMs), …

Temporal Action LocalizationAction Classification

Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere

2025-05-16 · Li Ju, Max Andersson, Stina Fredriksson, Edward Glöckner 외

Vision-language models (VLMs) as foundation models have significantly enhanced performance across a wide range of visual and textual tasks, without requiring large-scale training from scratch for downstream tasks. Howeve…

Uncertainty Quantification

Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation

2023-09-24 · NeurIPS 2023 11 · Yun Xing, Jian Kang, Aoran Xiao, Jiahao Nie 외

Vision-Language Pre-training has demonstrated its remarkable zero-shot recognition ability and potential to learn generalizable visual representations from language supervision. Taking a step ahead, language-supervised s…

SegmentationSemantic SegmentationZero-Shot Learning

Harnessing Large Language Models for Training-free Video Anomaly Detection

2024-04-01 · CVPR 2024 1 · Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang 외

Video anomaly detection (VAD) aims to temporally locate abnormal events in a video. Existing works mostly rely on training deep models to learn the distribution of normality with either video-level supervision, one-class…

Anomaly DetectionVideo Anomaly Detection