paper-with-me

홈 › Papers

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

2025-09-11 · Tianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng, Kaicheng Yang, Qichuan Ding arxiv

Although Contrastive Language-Image Pre-training (CLIP) exhibits strong performance across diverse vision tasks, its application to person representation learning faces two critical challenges: (i) the scarcity of large-scale annotated vision-language data focused on person-centric images, and (ii) the inherent limitations of global contrastive learning, which struggles to maintain discriminative local features crucial for fine-grained matching while remaining vulnerable to noisy text tokens. This work advances CLIP for person representation learning through synergistic improvements in data curation and model architecture. First, we develop a noise-resistant data construction pipeline that leverages the in-context learning capabilities of MLLMs to automatically filter and caption web-sourced images. This yields WebPerson, a large-scale dataset of 5M high-quality person-centric image-text pairs. Second, we introduce the GA-DMS (Gradient-Attention Guided Dual-Masking Synergetic) framework, which improves cross-modal alignment by adaptively masking noisy textual tokens based on the gradient-attention similarity score. Additionally, we incorporate masked token prediction objectives that compel the model to predict informative text tokens, enhancing fine-grained semantic representation learning. Extensive experiments show that GA-DMS achieves state-of-the-art performance across multiple benchmarks.

📄 PDF Abstract BibTeX arXiv:2509.09118

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningContrastive LearningPerson Retrieval

Similar Papers 제목 키워드 기반

How Do Decoder-Only LLMs Perceive Users? Rethinking Attention Masking for User Representation Learning

2026-02-11 · Jiahao Yuan, Yike Xu, Jinyong Wen, Baokun Wang 외 arxiv

Decoder-only large language models are increasingly used as behavioral encoders for user representation learning, yet the impact of attention masking on the quality of user embeddings remains underexplored. In this work,…

Representation LearningContrastive Learning

Semi-adaptive Synergetic Two-way Pseudoinverse Learning System

2024-06-27 · Binghong Liu, Ziqi Zhao, Shupan Li, Ke Wang

Deep learning has become a crucial technology for making breakthroughs in many fields. Nevertheless, it still faces two important challenges in theoretical and applied aspects. The first lies in the shortcomings of gradi…

HarmoGS: Robust 3D Gaussian Splatting in the Wild via Conflict-Aware Gradient Harmonization

2026-05-13 · Yulei Kang, Tianze Zhu, Jian-Fang Hu, Jianhuang Lai 외 arxiv

In-the-wild 3D Gaussian Splatting remains challenging due to transient distractors and illumination-induced cross-view appearance inconsistencies. Existing methods mainly rely on image-level masking to suppress unreliabl…

Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation

2025-02-26 · Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Domae Yukiyasu

This paper introduces a novel pipeline to enhance the precision of object masking for robotic manipulation within the specific domain of masking products in convenience stores. The approach integrates two advanced AI mod…

SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

2022-06-21 · Gang Li, Heliang Zheng, Daqing Liu, Chaoyue Wang 외

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (M…

Language ModelingLanguage ModellingMasked Language ModelingSemantic Segmentation