paper-with-me

홈 › Papers

Are All Combinations Equal? Combining Textual and Visual Features with Multiple Space Learning for Text-Based Video Retrieval

2022-11-21 · Damianos Galanopoulos, Vasileios Mezaris

In this paper we tackle the cross-modal video retrieval problem and, more specifically, we focus on text-to-video retrieval. We investigate how to optimally combine multiple diverse textual and visual features into feature pairs that lead to generating multiple joint feature spaces, which encode text-video pairs into comparable representations. To learn these representations our proposed network architecture is trained by following a multiple space learning procedure. Moreover, at the retrieval stage, we introduce additional softmax operations for revising the inferred query-video similarities. Extensive experiments in several setups based on three large-scale datasets (IACC.3, V3C1, and MSR-VTT) lead to conclusions on how to best combine text-visual features and document the performance of the proposed network. Source code is made publicly available at: https://github.com/bmezaris/TextToVideoRetrieval-TtimesV

📄 PDF Abstract BibTeX arXiv:2211.11351

Code (1)

bmezaris/texttovideoretrieval-ttimesv 공식 구현 pytorch

Tasks

AllRetrievalText to Video RetrievalVideo Retrieval

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Double Attention-based Multimodal Neural Machine Translation with Semantic Image Regions

2020-11-03 · EAMT 2020 11 · YuTing Zhao, Mamoru Komachi, Tomoyuki Kajiwara, Chenhui Chu

Existing studies on multimodal neural machine translation (MNMT) have mainly focused on the effect of combining visual and textual modalities to improve translations. However, it has been suggested that the visual modali…

Machine TranslationTranslation

A Probabilistic Model for Joint Learning of Word Embeddings from Texts and Images

2018-10-01 · EMNLP 2018 10 · Melissa Ailem, Bo-Wen Zhang, Aurelien Bellet, Pascal Denis 외

Several recent studies have shown the benefits of combining language and perception to infer word embeddings. These multimodal approaches either simply combine pre-trained textual and visual representations (e.g. feature…

Coreference ResolutionImage ClassificationQuestion AnsweringRetrieval+3

Combining Visual and Textual Features for Information Extraction from Online Flyers

2014-10-01 · EMNLP 2014 10 · Emilia Apostolova, Noriko Tomuro
Named Entity Recognition (NER)

Combining Multiple Views for Visual Speech Recognition

2017-10-19 · Marina Zimmermann, Mostafa Mehdipour Ghazi, Hazim Kemal Ekenel, Jean-Philippe Thiran

Visual speech recognition is a challenging research problem with a particular practical application of aiding audio speech recognition in noisy scenarios. Multiple camera setups can be beneficial for the visual speech re…

Sentencespeech-recognitionSpeech RecognitionVisual Speech Recognition

Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions

2015-09-01 · EMNLP 2015 9 · Arnau Ramisa, Josiah Wang, Ying Lu, Dell 외
Image Retrieval