paper-with-me

Papers

Deep Cross-Modal Projection Learning for Image-Text Matching

2018-09-01 · ECCV 2018 9 · Ying Zhang, Huchuan Lu

The key point of image-text matching is how to accurately measure the similarity between visual and textual inputs. Despite the great progress of associating the deep cross-modal embeddings with the bi-directional ranking loss, developing the strategies for mining useful triplets and selecting appropriate margins remains a challenge in real applications. In this paper, we propose a cross-modal projection matching (CMPM) loss and a cross-modal projection classification (CMPC) loss for learning discriminative image-text embeddings. The CMPM loss minimizes the KL divergence between the projection compatibility distributions and the normalized matching distributions defined with all the positive and negative samples in a mini-batch. The CMPC loss attempts to categorize the vector projection of representations from one modality onto another with the improved norm-softmax loss, for further enhancing the feature compactness of each class. Extensive analysis and experiments on multiple datasets demonstrate the superiority of the proposed approach.

📄 PDF Abstract BibTeX

Code (1)

YingZhangDUT/Cross-Modal-Projection-Learning 공식 구현 tf

Tasks

Cross-Modal RetrievalImage-text matchingText based Person RetrievalText Matching

Similar Papers 제목 키워드 기반

Dual-path CNN with Max Gated block for Text-Based Person Re-identification

2020-09-20 · Tinghuai Ma, Mingming Yang, Huan Rong, Yurong Qian 외

Text-based person re-identification(Re-id) is an important task in video surveillance, which consists of retrieving the corresponding person's image given a textual description from a large gallery of images. It is diffi…

Language ModellingPerson Re-IdentificationWord Embeddings

A Concept-Centric Approach to Multi-Modality Learning

2024-12-18 · Yuchong Geng, Ao Tang

In an effort to create a more efficient AI system, we introduce a new multi-modality learning framework that leverages a modality-agnostic concept space possessing abstract knowledge and a set of modality-specific projec…

Image-text matchingQuestion AnsweringText MatchingVisual Question Answering

Weakly Supervised Text-Based Person Re-Identification

2021-01-01 · ICCV 2021 10 · Shizhen Zhao, Changxin Gao, Yuanjie Shao, Wei-Shi Zheng 외

The conventional text-based person re-identification methods heavily rely on identity annotations. However, this labeling process is costly and time-consuming. In this paper, we consider a more practical setting call…

ClusteringPerson Re-IdentificationPseudo Label

Image-based Geolocalization by Ground-to-2.5D Map Matching

2023-08-11 · Mengjie Zhou, Liu Liu, Yiran Zhong, Andrew Calway

We study the image-based geolocalization problem, aiming to localize ground-view query images on cartographic maps. Current methods often utilize cross-view localization techniques to match ground-view query images with …

Image-Based Localization

Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation

2025-05-16 · Xin Zhang, Ziruo Zhang, Jiawei Du, Zuozhu Liu 외

Multimodal Dataset Distillation (MDD) seeks to condense large-scale image-text datasets into compact surrogates while retaining their effectiveness for cross-modal learning. Despite recent progress, existing MDD approach…

cross-modal alignmentDataset Distillation