paper-with-me

홈 › Papers

Prototype-guided Cross-modal Completion and Alignment for Incomplete Text-based Person Re-identification

2023-09-29 · Tiantian Gong, Guodong Du, Junsheng Wang, Yongkang Ding, Liyan Zhang

Traditional text-based person re-identification (ReID) techniques heavily rely on fully matched multi-modal data, which is an ideal scenario. However, due to inevitable data missing and corruption during the collection and processing of cross-modal data, the incomplete data issue is usually met in real-world applications. Therefore, we consider a more practical task termed the incomplete text-based ReID task, where person images and text descriptions are not completely matched and contain partially missing modality data. To this end, we propose a novel Prototype-guided Cross-modal Completion and Alignment (PCCA) framework to handle the aforementioned issues for incomplete text-based ReID. Specifically, we cannot directly retrieve person images based on a text query on missing modality data. Therefore, we propose the cross-modal nearest neighbor construction strategy for missing data by computing the cross-modal similarity between existing images and texts, which provides key guidance for the completion of missing modal features. Furthermore, to efficiently complete the missing modal features, we construct the relation graphs with the aforementioned cross-modal nearest neighbor sets of missing modal data and the corresponding prototypes, which can further enhance the generated missing modal features. Additionally, for tighter fine-grained alignment between images and texts, we raise a prototype-aware cross-modal alignment loss that can effectively reduce the modality heterogeneity gap for better fine-grained alignment in common space. Extensive experimental results on several benchmarks with different missing ratios amply demonstrate that our method can consistently outperform state-of-the-art text-image ReID approaches.

📄 PDF Abstract BibTeX arXiv:2309.17104

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignmentPerson Re-Identification

Similar Papers 제목 키워드 기반

Explicitly Guided Information Interaction Network for Cross-modal Point Cloud Completion

2024-07-03 · Hang Xu, Chen Long, Wenxiao Zhang, YuAn Liu 외

In this paper, we explore a novel framework, EGIInet (Explicitly Guided Information Interaction Network), a model for View-guided Point cloud Completion (ViPC) task, which aims to restore a complete point cloud from a pa…

Point Cloud Completion

MAGE: View-guided Point Cloud Completion with Efficient Modality Alignment and Adaptive Geometry Enhancement

2026-06-30 · Weize Quan, Zhengwei Wu, Kai Wang, Dong-Ming Yan arxiv

View-based point cloud completion aims to recover a complete 3D shape from a partial point cloud, guided by a single-view image. However, existing approaches often suffer from limited performance due to weak modality ali…

Point Cloud Completion

Disentangled Fine-Grained Prototype Learning for Incomplete Image-Tabular Classification

2026-06-03 · Feixiang Zhou, Jianyang Xie, Zhuangzhi Gao, Qinkai Yu 외 arxiv

The missing-modality problem poses a significant challenge in image-tabular multimodal learning across a wide range of multimedia applications, including product understanding, recommendation systems, and medical diagnos…

Recommendation SystemsMedical Diagnosis

MVCL-DAF++: Enhancing Multimodal Intent Recognition via Prototype-Aware Contrastive Alignment and Coarse-to-Fine Dynamic Attention Fusion

2025-09-22 · Haofeng Huang, Yifei Han, Long Zhang, Bin Li 외 arxiv

Multimodal intent recognition (MMIR) suffers from weak semantic grounding and poor robustness under noisy or rare-class conditions. We propose MVCL-DAF++, which extends MVCL-DAF with two key modules: (1) Prototype-aware …

Multimodal Intent Recognition

ProtoMol: Enhancing Molecular Property Prediction via Prototype-Guided Multimodal Learning

2025-10-19 · Yingxu Wang, Kunyu Zhang, Jiaxin Huang, Nan Yin 외 arxiv

Multimodal molecular representation learning, which jointly models molecular graphs and their textual descriptions, enhances predictive accuracy and interpretability by enabling more robust and reliable predictions of dr…

Molecular Property PredictionRepresentation Learning