Support-set based Multi-modal Representation Enhancement for Video Captioning
Video captioning is a challenging task that necessitates a thorough comprehension of visual scenes. Existing methods follow a typical one-to-one mapping, which concentrates on a limited sample space while ignoring the intrinsic semantic associations between samples, resulting in rigid and uninformative expressions. To address this issue, we propose a novel and flexible framework, namely Support-set based Multi-modal Representation Enhancement (SMRE) model, to mine rich information in a semantic subspace shared between samples. Specifically, we propose a Support-set Construction (SC) module to construct a support-set to learn underlying connections between samples and obtain semantic-related visual elements. During this process, we design a Semantic Space Transformation (SST) module to constrain relative distance and administrate multi-modal interactions in a self-supervised way. Extensive experiments on MSVD and MSR-VTT datasets demonstrate that our SMRE achieves state-of-the-art performance.
Code (1)
Tasks
Video CaptioningSimilar Papers 제목 키워드 기반
AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference
Low-light video enhancement (LLVE) remains a challenging task due to severe information degradation under low-illumination conditions. Recent multimodal approaches have significantly improved enhancement performance by i…
Video EnhancementTencent Text-Video Retrieval: Hierarchical Cross-Modal Interactions with Multi-Level Representations
Text-Video Retrieval plays an important role in multi-modal understanding and has attracted increasing attention in recent years. Most existing methods focus on constructing contrastive pairs between whole videos and com…
Contrastive LearningDenoisingRetrievalSentence+2MAiVAR-T: Multimodal Audio-image and Video Action Recognizer using Transformers
In line with the human capacity to perceive the world by simultaneously processing and integrating high-dimensional inputs from multiple modalities like vision and audio, we propose a novel model, MAiVAR-T (Multimodal Au…
Action RecognitionTemporal Action LocalizationRepresentation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues …
Action RecognitionDenoisingRepresentation LearningTemporal Action LocalizationTECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction
The objective of multimodal intent recognition (MIR) is to leverage various modalities-such as text, video, and audio-to detect user intentions, which is crucial for understanding human language and context in dialogue s…
Intent RecognitionMultimodal Intent Recognition