paper-with-me

홈 › Papers

Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text Retrieval

2018-06-11 · ICMR 2018 6 · Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, Amit K. Roy-Chowdhury

Constructing a joint representation invariant across different modalities (e.g., video, language) is of significant importance in many multimedia applications. While there are a number of recent successes in developing effective image-text retrieval methods by learning joint representations, the video-text retrieval task, in contrast, has not been explored to its fullest extent. In this paper, we study how to effectively utilize available multi-modal cues from videos for the cross-modal video-text retrieval task. Based on our analysis, we propose a novel framework that simultaneously utilizes multimodal features (different visual characteristics, audio inputs, and text) by a fusion strategy for efficient retrieval. Furthermore, we explore several loss functions in training the joint embedding and propose a modified pairwise ranking loss for the retrieval task. Experiments on MSVD and MSR-VTT datasets demonstrate that our method achieves significant performance gain compared to the state-of-the-art approaches.

📄 PDF Abstract BibTeX

Code (1)

niluthpol/multimodal_vtt pytorch

Tasks

Image-text RetrievalRetrievalText RetrievalVideo RetrievalVideo-Text Retrieval

Similar Papers 제목 키워드 기반

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

2026-05-17 · Yuheng Chen, Qingdong He, Teng Hu, Yuji Wang 외 arxiv

The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides, achieving cohesive multimodal customization for the simultaneous pr…

Video Generation

Uncovering Grounding IDs: How External Cues Shape Multimodal Binding

2025-09-28 · Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari, Mobin Bagherian 외 arxiv

Large vision-language models (LVLMs) show strong performance across multimodal benchmarks but remain limited in structured reasoning and precise grounding. Recent work has demonstrated that adding simple visual structure…

CrossFlowDG: Bridging the Modality Gap with Cross-modal Flow Matching for Domain Generalization

2026-04-18 · Antonios Kritikos, Nikolaos Spanos, Athanasios Voulodimos arxiv

Domain generalization (DG) aims to maintain performance under domain shift, which in computer vision appears primarily as stylistic variations that cause models to overfit to domain-specific appearance cues rather than c…

Semantic correspondenceDomain Generalization

Learning Social Image Embedding with Deep Multimodal Attention Networks

2017-10-18 · Feiran Huang, Xiao-Ming Zhang, Zhoujun Li, Tao Mei 외

Learning social media data embedding by deep models has attracted extensive research interest as well as boomed a lot of applications, such as link prediction, classification, and cross-modal search. However, for social …

ClassificationGeneral ClassificationLink PredictionMulti-Label Classification+2

Tracing Intricate Cues in Dialogue: Joint Graph Structure and Sentiment Dynamics for Multimodal Emotion Recognition

2024-07-31 · Jiang Li, XiaoPing Wang, Zhigang Zeng

Multimodal emotion recognition in conversation (MERC) has garnered substantial research attention recently. Existing MERC methods face several challenges: (1) they fail to fully harness direct inter-modal cues, possibly …

Emotion RecognitionEmotion Recognition in ConversationMultimodal Emotion RecognitionMultimodal Sentiment Analysis+1