paper-with-me

Papers

Support-set based Multi-modal Representation Enhancement for Video Captioning

2022-05-19 · Xiaoya Chen, Jingkuan Song, Pengpeng Zeng, Lianli Gao, Heng Tao Shen

Video captioning is a challenging task that necessitates a thorough comprehension of visual scenes. Existing methods follow a typical one-to-one mapping, which concentrates on a limited sample space while ignoring the intrinsic semantic associations between samples, resulting in rigid and uninformative expressions. To address this issue, we propose a novel and flexible framework, namely Support-set based Multi-modal Representation Enhancement (SMRE) model, to mine rich information in a semantic subspace shared between samples. Specifically, we propose a Support-set Construction (SC) module to construct a support-set to learn underlying connections between samples and obtain semantic-related visual elements. During this process, we design a Semantic Space Transformation (SST) module to constrain relative distance and administrate multi-modal interactions in a self-supervised way. Extensive experiments on MSVD and MSR-VTT datasets demonstrate that our SMRE achieves state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2205.09307

Code (1)

smre-cv/smre 공식 구현 pytorch

Tasks

Video Captioning

Similar Papers 제목 키워드 기반

AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference

2026-06-09 · Hangfeng Liang, Yutao Hu, Yanhan Hu, Xiaohan Wu 외 arxiv

Low-light video enhancement (LLVE) remains a challenging task due to severe information degradation under low-illumination conditions. Recent multimodal approaches have significantly improved enhancement performance by i…

Video Enhancement

Tencent Text-Video Retrieval: Hierarchical Cross-Modal Interactions with Multi-Level Representations

2022-04-07 · Jie Jiang, Shaobo Min, Weijie Kong, Dihong Gong 외

Text-Video Retrieval plays an important role in multi-modal understanding and has attracted increasing attention in recent years. Most existing methods focus on constructing contrastive pairs between whole videos and com…

Contrastive LearningDenoisingRetrievalSentence+2

MAiVAR-T: Multimodal Audio-image and Video Action Recognizer using Transformers

2023-08-01 · Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, Naveed Akhtar

In line with the human capacity to perceive the world by simultaneously processing and integrating high-dimensional inputs from multiple modalities like vision and audio, we propose a novel model, MAiVAR-T (Multimodal Au…

Action RecognitionTemporal Action Localization

Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement

2022-05-07 · Bing Li, Jiaxin Chen, Dongming Zhang, Xiuguo Bao 외

Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues …

Action RecognitionDenoisingRepresentation LearningTemporal Action Localization

TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction

2024-12-11 · Quynh-Mai Thi Nguyen, Lan-Nhi Thi Nguyen, Cam-Van Thi Nguyen

The objective of multimodal intent recognition (MIR) is to leverage various modalities-such as text, video, and audio-to detect user intentions, which is crucial for understanding human language and context in dialogue s…

Intent RecognitionMultimodal Intent Recognition