paper-with-me

홈 › Papers

Cross-modal Deep Metric Learning with Multi-task Regularization

2017-03-21 · Xin Huang, Yuxin Peng

DNN-based cross-modal retrieval has become a research hotspot, by which users can search results across various modalities like image and text. However, existing methods mainly focus on the pairwise correlation and reconstruction error of labeled data. They ignore the semantically similar and dissimilar constraints between different modalities, and cannot take advantage of unlabeled data. This paper proposes Cross-modal Deep Metric Learning with Multi-task Regularization (CDMLMR), which integrates quadruplet ranking loss and semi-supervised contrastive loss for modeling cross-modal semantic similarity in a unified multi-task learning architecture. The quadruplet ranking loss can model the semantically similar and dissimilar constraints to preserve cross-modal relative similarity ranking information. The semi-supervised contrastive loss is able to maximize the semantic similarity on both labeled and unlabeled data. Compared to the existing methods, CDMLMR exploits not only the similarity ranking information but also unlabeled cross-modal data, and thus boosts cross-modal retrieval accuracy.

📄 PDF Abstract BibTeX arXiv:1703.07026

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalMetric LearningMulti-Task LearningRetrievalSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Unimodal Cyclic Regularization for Training Multimodal Image Registration Networks

2020-11-12 · Zhe Xu, Jiangpeng Yan, Jie Luo, William Wells 외

The loss function of an unsupervised multimodal image registration framework has two terms, i.e., a metric for similarity measure and regularization. In the deep learning era, researchers proposed many approaches to auto…

Image Registration

MOVER: Multimodal Optimal Transport with Volume-based Embedding Regularization

2025-08-16 · Haochen You, Baojing Liu arxiv

Recent advances in multimodal learning have largely relied on pairwise contrastive objectives to align different modalities, such as text, video, and audio, in a shared embedding space. While effective in bi-modal setups…

Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion

2026-01-29 · Zixuan Xia, Hao Wang, Pengcheng Weng, Yanyu Qian 외 arxiv

Multimodal fusion is often treated as an optimization-balancing problem, where training signals are adjusted to prevent one modality from dominating the others. However, balanced optimization does not fully determine the…

Representation Learning

Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task

2025-09-04 · JungHo Jung, Junhyun Lee arxiv

End-to-end speech-to-text translation typically suffers from the scarcity of paired speech-text data. One way to overcome this shortcoming is to utilize the bitext data from the Machine Translation (MT) task and perform …

Speech-to-Text TranslationMachine TranslationMulti-Task Learning

Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning

2023-03-10 · CVPR 2023 1 · Qian Jiang, Changyou Chen, Han Zhao, Liqun Chen 외

Contrastive loss has been increasingly used in learning representations from multiple modalities. In the limit, the nature of the contrastive loss encourages modalities to exactly match each other in the latent space. Ye…

Few-Shot Image Classificationimage-classificationImage ClassificationImage-text Retrieval+9