paper-with-me

홈 › Papers

Learning Shared Semantic Space with Correlation Alignment for Cross-modal Event Retrieval

2019-01-14 · Zhenguo Yang, Zehang Lin, Peipei Kang, Jianming Lv, Qing Li, Wenyin Liu

In this paper, we propose to learn shared semantic space with correlation alignment (${S}^{3}CA$) for multimodal data representations, which aligns nonlinear correlations of multimodal data distributions in deep neural networks designed for heterogeneous data. In the context of cross-modal (event) retrieval, we design a neural network with convolutional layers and fully-connected layers to extract features for images, including images on Flickr-like social media. Simultaneously, we exploit a fully-connected neural network to extract semantic features for texts, including news articles from news media. In particular, nonlinear correlations of layer activations in the two neural networks are aligned with correlation alignment during the joint training of the networks. Furthermore, we project the multimodal data into a shared semantic space for cross-modal (event) retrieval, where the distances between heterogeneous data samples can be measured directly. In addition, we contribute a Wiki-Flickr Event dataset, where the multimodal data samples are not describing each other in pairs like the existing paired datasets, but all of them are describing semantic events. Extensive experiments conducted on both paired and unpaired datasets manifest the effectiveness of ${S}^{3}CA$, outperforming the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:1901.04268

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesRetrieval

Similar Papers 제목 키워드 기반

Dual-Stream Cross-Modal Representation Learning via Residual Semantic Decorrelation

2025-12-08 · Xuecheng Li, Weikuan Jia, Alisher Kurbonaliev, Qurbonaliev Alisher 외 arxiv

Cross-modal learning has become a fundamental paradigm for integrating heterogeneous information sources such as images, text, and structured attributes. However, multimodal representations often suffer from modality dom…

Representation Learning

Multi-lingual Common Semantic Space Construction via Cluster-consistent Word Embedding

2018-04-21 · EMNLP 2018 10 · Lifu Huang, Kyunghyun Cho, Boliang Zhang, Heng Ji 외

We construct a multilingual common semantic space based on distributional semantics, where words from multiple languages are projected into a shared space to enable knowledge and resource transfer across languages. Beyon…

ClusteringWord Alignment

Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

2026-06-28 · Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon hf

Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute bud…

Cross-Modal Retrieval

Disentangled Fine-Grained Prototype Learning for Incomplete Image-Tabular Classification

2026-06-03 · Feixiang Zhou, Jianyang Xie, Zhuangzhi Gao, Qinkai Yu 외 arxiv

The missing-modality problem poses a significant challenge in image-tabular multimodal learning across a wide range of multimedia applications, including product understanding, recommendation systems, and medical diagnos…

Recommendation SystemsMedical Diagnosis

LatentUMM: Dual Latent Alignment for Unified Multimodal Models

2026-05-18 · Yinyi Luo, Wenwen Wang, Hayes Bai, Marios Savvides 외 arxiv

Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We obser…