paper-with-me

홈 › Papers

Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition

2025-01-01 · CVPR 2025 1 · Junyi Wu, Yan Huang, Min Gao, Yuzhen Niu, Yuzhong Chen, Qiang Wu

Pedestrian attribute recognition (PAR) seeks to predict multiple semantic attributes associated with a specific pedestrian. There are two types of approaches for PAR: unimodal framework and bimodal framework. The former one is to seek a robust visual feature. However, the lack of exploiting semantic feature of linguistic modality is the main concern. The latter one utilizes prompt learning techniques to integrate linguistic data. However, static prompt templates and simple bimodal concatenation cannot to capture the extensive intra-class attribute variability and support active modalities collaboration. In this paper, we propose an Enhanced Visual-Semantic Interaction with Tailored Prompts (EVSITP) framework for PAR. We present an Image-Conditional Dual-Prompt Initialization Module (IDIM) to adaptively generate context-sensitive prompts from visual inputs. Subsequently, a Prompt Enhanced and Regularization Module (PERM) is proposed to strengthen linguistic information from IDIM. We further design a Bimodal Mutual Interaction Module (BMIM) to ensure bidirectional modalities communication. In addition, existing PAR datasets are collected over a short period in limited scenarios, which do not align with real-world scenarios. Therefore, we annotate a long-term person re-identification dataset to create a new PAR dataset, Celeb-PAR. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

AttributePedestrian Attribute RecognitionPerson Re-IdentificationPrompt Learning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Tailored Visions: Enhancing Text-to-Image Generation with Personalized Prompt Rewriting

2023-10-12 · CVPR 2024 1 · Zijie Chen, Lichao Zhang, Fangsheng Weng, Lili Pan 외

Despite significant progress in the field, it is still challenging to create personalized visual representations that align closely with the desires and preferences of individual users. This process requires users to art…

Image GenerationText to Image GenerationText-to-Image Generation

Vision-Language Enhanced Foundation Model for Semi-supervised Medical Image Segmentation

2025-11-24 · Jiaqi Guo, Mingzhen Li, Hanyu Su, Santiago López 외 arxiv

Semi-supervised learning (SSL) has emerged as an effective paradigm for medical image segmentation, reducing the reliance on extensive expert annotations. Meanwhile, vision-language models (VLMs) have demonstrated strong…

Semi-supervised Medical Image Segmentation

KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration

2025-01-07 · Chengyuan Li, Suyang Zhou, Jieping Kong, Lei Qi 외

Zero-shot anomaly detection (ZSAD) identifies anomalies without needing training samples from the target dataset, essential for scenarios with privacy concerns or limited data. Vision-language models like CLIP show poten…

Anomaly DetectionAnomaly SegmentationGeneral KnowledgeLarge Language Model+4

Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors

2024-05-17 · Jiachen Sun, Changsheng Wang, Jiongxiao Wang, Yiwei Zhang 외

Large language models have become increasingly prominent, also signaling a shift towards multimodality as the next frontier in artificial intelligence, where their embeddings are harnessed as prompts to generate textual …

Adversarial Attack

KNN Transformer with Pyramid Prompts for Few-Shot Learning

2024-10-14 · Wenhao Li, Qiangchang Wang, Peng Zhao, Yilong Yin

Few-Shot Learning (FSL) aims to recognize new classes with limited labeled data. Recent studies have attempted to address the challenge of rare samples with textual prompts to modulate visual features. However, they usua…

Few-Shot Learning