paper-with-me

Papers

Multiple Visual-Semantic Embedding for Video Retrieval from Query Sentence

2020-04-16 · Huy Manh Nguyen, Tomo Miyazaki, Yoshihiro Sugaya, Shinichiro Omachi

Visual-semantic embedding aims to learn a joint embedding space where related video and sentence instances are located close to each other. Most existing methods put instances in a single embedding space. However, they struggle to embed instances due to the difficulty of matching visual dynamics in videos to textual features in sentences. A single space is not enough to accommodate various videos and sentences. In this paper, we propose a novel framework that maps instances into multiple individual embedding spaces so that we can capture multiple relationships between instances, leading to compelling video retrieval. We propose to produce a final similarity between instances by fusing similarities measured in each embedding space using a weighted sum strategy. We determine the weights according to a sentence. Therefore, we can flexibly emphasize an embedding space. We conducted sentence-to-video retrieval experiments on a benchmark dataset. The proposed method achieved superior performance, and the results are competitive to state-of-the-art methods. These experimental results demonstrated the effectiveness of the proposed multiple embedding approach compared to existing methods.

📄 PDF Abstract BibTeX arXiv:2004.07967

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalSentenceVideo Retrieval

Similar Papers 제목 키워드 기반

Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval

2019-06-11 · CVPR 2019 6 · Yale Song, Mohammad Soleymani

Visual-semantic embedding aims to find a shared latent space where related visual and textual instances are close to each other. Most current methods learn injective embedding functions that map an instance to a single p…

Cross-Modal RetrievalMultiple Instance LearningRetrievalSentence+2

GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval

2025-08-03 · Bowen Yang, Yun Cao, Chen He, Xiaosu Su arxiv

Text-to-video retrieval requires precise alignment between language and temporally rich audio-video signals. However, existing methods often emphasize visual cues while underutilizing audio semantics or relying on coarse…

Computational EfficiencyVideo Retrieval

Query by Activity Video in the Wild

2023-11-23 · Tao Hu, William Thong, Pascal Mettes, Cees G. M. Snoek

This paper focuses on activity retrieval from a video query in an imbalanced scenario. In current query-by-activity-video literature, a common assumption is that all activities have sufficient labelled examples when lear…

AllRetrieval

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

2025-07-07 · Rui Meng, Ziyan Jiang, Ye Liu, Mingyi Su 외 arxiv

Multimodal embedding models have been crucial in enabling various downstream tasks such as semantic similarity, information retrieval, and clustering over different modalities. However, existing multimodal embeddings lik…

Video Question AnsweringRepresentation LearningInformation RetrievalVideo Classification

MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation

2025-12-15 · Huu-An Vu, Van-Khanh Mai, Trong-Tam Nguyen, Quang-Duc Dam 외 arxiv

The rapid expansion of video content across online platforms has accelerated the need for retrieval systems capable of understanding not only isolated visual moments but also the temporal structure of complex events. Exi…

Visual GroundingVideo Retrieval