paper-with-me

홈 › Papers

Integrating multi-label contrastive learning with dual adversarial graph neural networks for cross-modal retrieval

2022-07-05 · IEEE Transactions on Pattern Analysis and Machine Intelligence 2022 7 · Dizhan Xue, Shengsheng Qian, Quan Fang, Changsheng Xu

With the growing amount of multimodal data, cross-modal retrieval has attracted more and more attention and become a hot research topic. To date, most of the existing techniques mainly convert multimodal data into a common representation space where similarities in semantics between samples can be easily measured across multiple modalities. However, these approaches may suffer from the following limitations: 1) They overcome the modality gap by introducing loss in the common representation space, which may not be sufficient to eliminate the heterogeneity of various modalities; 2) They treat labels as independent entities and ignore label relationships, which is not conducive to establishing semantic connections across multimodal data; 3) They ignore the non-binary values of label similarity in multi-label scenarios, which may lead to inefficient alignment of representation similarity with label similarity. To tackle these problems, in this article, we propose two models to learn discriminative and modality-invariant representations for cross-modal retrieval. First, the dual generative adversarial networks are built to project multimodal data into a common representation space. Second, to model label relation dependencies and develop inter-dependent classifiers, we employ multi-hop graph neural networks (consisting of Probabilistic GNN and Iterative GNN), where the layer aggregation mechanism is suggested for using propagation information of various hops. Third, we propose a novel soft multi-label contrastive loss for cross-modal retrieval, with the soft positive sampling probability, which can align the representation similarity and the label similarity. Additionally, to adapt to incomplete-modal learning, which can have wider applications, we propose a modal reconstruction mechanism to generate missing features. Extensive experiments on three widely used benchmark datasets, i.e., NUS-WIDE, MIRFlickr, and MS-COCO, show the superiority of our proposed method.

📄 PDF Abstract BibTeX

Code (1)

LivXue/GNN4CMR pytorch

Tasks

Contrastive LearningCross-Modal RetrievalRetrieval

Similar Papers 제목 키워드 기반

Decoupled Adversarial Contrastive Learning for Self-supervised Adversarial Robustness

2022-07-22 · Chaoning Zhang, Kang Zhang, Chenshuang Zhang, Axi Niu 외

Adversarial training (AT) for robust representation learning and self-supervised learning (SSL) for unsupervised representation learning are two active research fields. Integrating AT into SSL, multiple prior works have …

Adversarial RobustnessContrastive LearningPhilosophyRepresentation Learning+1

CATformer: Contrastive Adversarial Transformer for Image Super-Resolution

2025-08-25 · Qinyi Tian, Spence Cox, Laura E. Dalton arxiv

Super-resolution remains a promising technique to enhance the quality of low-resolution images. This study introduces CATformer (Contrastive Adversarial Transformer), a novel neural network integrating diffusion-inspired…

Image Super-ResolutionContrastive Learning

Weakly Supervised Contrastive Adversarial Training for Learning Robust Features from Semi-supervised Data

2025-03-14 · CVPR 2025 1 · Lilin Zhang, Chengpei Wu, Ning Yang

Existing adversarial training (AT) methods often suffer from incomplete perturbation, meaning that not all non-robust features are perturbed when generating adversarial examples (AEs). This results in residual correlatio…

Semi-Supervised Dual-Stream Self-Attentive Adversarial Graph Contrastive Learning for Cross-Subject EEG-based Emotion Recognition

2023-08-13 · Weishan Ye, Zhiguo Zhang, Fei Teng, Min Zhang 외

Electroencephalography (EEG) is an objective tool for emotion recognition with promising applications. However, the scarcity of labeled data remains a major challenge in this field, limiting the widespread use of EEG-bas…

Contrastive LearningDomain AdaptationEEGEmotion Recognition

SLADE: Shielding against Dual Exploits in Large Vision-Language Models

2025-01-01 · CVPR 2025 1 · Md Zarif Hossain, Ahmed Imteaj

Large Vision-Language Models (LVLMs) have emerged as transformative tools in multimodal tasks, seamlessly integrating pretrained vision encoders to align visual and textual modalities. Prior works have highlighted th…

Contrastive LearningInstruction Following