paper-with-me

Papers

CM-GANs: Cross-modal Generative Adversarial Networks for Common Representation Learning

2017-10-14 · Yuxin Peng, Jinwei Qi, Yuxin Yuan

It is known that the inconsistent distribution and representation of different modalities, such as image and text, cause the heterogeneity gap that makes it challenging to correlate such heterogeneous data. Generative adversarial networks (GANs) have shown its strong ability of modeling data distribution and learning discriminative representation, existing GANs-based works mainly focus on generative problem to generate new data. We have different goal, aim to correlate heterogeneous data, by utilizing the power of GANs to model cross-modal joint distribution. Thus, we propose Cross-modal GANs to learn discriminative common representation for bridging heterogeneity gap. The main contributions are: (1) Cross-modal GANs architecture is proposed to model joint distribution over data of different modalities. The inter-modality and intra-modality correlation can be explored simultaneously in generative and discriminative models. Both of them beat each other to promote cross-modal correlation learning. (2) Cross-modal convolutional autoencoders with weight-sharing constraint are proposed to form generative model. They can not only exploit cross-modal correlation for learning common representation, but also preserve reconstruction information for capturing semantic consistency within each modality. (3) Cross-modal adversarial mechanism is proposed, which utilizes two kinds of discriminative models to simultaneously conduct intra-modality and inter-modality discrimination. They can mutually boost to make common representation more discriminative by adversarial training process. To the best of our knowledge, our proposed CM-GANs approach is the first to utilize GANs to perform cross-modal common representation learning. Experiments are conducted to verify the performance of our proposed approach on cross-modal retrieval paradigm, compared with 10 methods on 3 cross-modal datasets.

📄 PDF Abstract BibTeX arXiv:1710.05106

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalRepresentation LearningRetrieval

Similar Papers 제목 키워드 기반

SyncGAN: Synchronize the Latent Space of Cross-modal Generative Adversarial Networks

2018-04-02 · Wen-Cheng Chen, Chien-Wen Chen, Min-Chun Hu

Generative adversarial network (GAN) has achieved impressive success on cross-domain generation, but it faces difficulty in cross-modal generation due to the lack of a common distribution between heterogeneous data. Most…

Generative Adversarial Network

Symbiotic Adversarial Learning for Attribute-based Person Search

2020-07-19 · ECCV 2020 8 · Yu-Tong Cao, Jingya Wang, DaCheng Tao

Attribute-based person search is in significant demand for applications where no detected query images are available, such as identifying a criminal from witness. However, the task itself is quite challenging because the…

Attributecross-modal alignmentPerson Search

On the Performance of Generative Adversarial Network (GAN) Variants: A Clinical Data Study

2020-09-21 · Jaesung Yoo, Jeman Park, An Wang, David Mohaisen 외

Generative Adversarial Network (GAN) is a useful type of Neural Networks in various types of applications including generative models and feature extraction. Various types of GANs are being researched with different insi…

Generative Adversarial Network

Emotion Detection Using Conditional Generative Adversarial Networks (cGAN): A Deep Learning Approach

2025-08-06 · Anushka Srivastava arxiv

This paper presents a deep learning-based approach to emotion detection using Conditional Generative Adversarial Networks (cGANs). Unlike traditional unimodal techniques that rely on a single data type, we explore a mult…

Emotion Recognition

Generative Adversarial Networks for Brain Images Synthesis: A Review

2023-05-16 · Firoozeh Shomal Zadeh, Sevda Molani, Maysam Orouskhani, Marziyeh Rezaei 외

In medical imaging, image synthesis is the estimation process of one image (sequence, modality) from another image (sequence, modality). Since images with different modalities provide diverse biomarkers and capture vario…

Deep LearningGenerative Adversarial NetworkImage Generation