paper-with-me

홈 › Papers

Dialogue-to-Video Retrieval

2023-03-23 · Chenyang Lyu, Manh-Duy Nguyen, Van-Tu Ninh, Liting Zhou, Cathal Gurrin, Jennifer Foster

Recent years have witnessed an increasing amount of dialogue/conversation on the web especially on social media. That inspires the development of dialogue-based retrieval, in which retrieving videos based on dialogue is of increasing interest for recommendation systems. Different from other video retrieval tasks, dialogue-to-video retrieval uses structured queries in the form of user-generated dialogue as the search descriptor. We present a novel dialogue-to-video retrieval system, incorporating structured conversational information. Experiments conducted on the AVSD dataset show that our proposed approach using plain-text queries improves over the previous counterpart model by 15.8% on R@1. Furthermore, our approach using dialogue as a query, improves retrieval performance by 4.2%, 6.2%, 8.6% on R@1, R@5 and R@10 and outperforms the state-of-the-art model by 0.7%, 3.6% and 6.0% on R@1, R@5 and R@10 respectively.

📄 PDF Abstract BibTeX arXiv:2303.16761

Code (1)

lyuchenyang/dialogue-to-video-retrieval 공식 구현 pytorch

Tasks

Recommendation SystemsRetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval

2026-02-22 · Diogo Glória-Silva, David Semedo, João Maglhães arxiv

We introduce VIGiA, a novel multimodal dialogue model designed to understand and reason over complex, multi-step instructional video action plans. Unlike prior work which focuses mainly on text-only guidance, or treats v…

VGNMN: Video-grounded Neural Module Networks for Video-Grounded Dialogue Systems

2022-07-01 · NAACL 2022 7 · Hung Le, Nancy Chen, Steven Hoi

Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded dialogue ta…

Information RetrievalQuestion AnsweringRetrievalVisual Question Answering+1

Game-Based Video-Context Dialogue

2018-09-12 · EMNLP 2018 10 · Ramakanth Pasunuru, Mohit Bansal

Current dialogue systems focus more on textual and speech context knowledge and are usually based on two speakers. Some recent work has investigated static image-based dialogue. However, several real-world human interact…

Retrieval

Simple Baselines for Interactive Video Retrieval with Questions and Answers

2023-08-21 · ICCV 2023 1 · Kaiqu Liang, Samuel Albanie

To date, the majority of video retrieval systems have been optimized for a "single-shot" scenario in which the user submits a query in isolation, ignoring previous interactions with the system. Recently, there has been r…

Question AnsweringRetrievalVideo Retrieval

VGNMN: Video-grounded Neural Module Network to Video-Grounded Language Tasks

2021-04-16 · Hung Le, Nancy F. Chen, Steven C. H. Hoi

Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded dialogue ta…

Information RetrievalQuestion AnsweringRetrievalVisual Question Answering+1