paper-with-me

홈 › Papers

Where to Play: Retrieval of Video Segments using Natural-Language Queries

2017-07-02 · Sangkuk Lee, Daesik Kim, Myunggi Lee, Jihye Hwang, Nojun Kwak

In this paper, we propose a new approach for retrieval of video segments using natural language queries. Unlike most previous approaches such as concept-based methods or rule-based structured models, the proposed method uses image captioning model to construct sentential queries for visual information. In detail, our approach exploits multiple captions generated by visual features in each image with `Densecap'. Then, the similarities between captions of adjacent images are calculated, which is used to track semantically similar captions over multiple frames. Besides introducing this novel idea of 'tracking by captioning', the proposed method is one of the first approaches that uses a language generation model learned by neural networks to construct semantic query describing the relations and properties of visual information. To evaluate the effectiveness of our approach, we have created a new evaluation dataset, which contains about 348 segments of scenes in 20 movie-trailers. Through quantitative and qualitative evaluation, we show that our method is effective for retrieval of video segments using natural language queries.

📄 PDF Abstract BibTeX arXiv:1707.00251

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningNatural Language QueriesRetrievalText Generation

Similar Papers 제목 키워드 기반

Large-Scale Query-by-Image Video Retrieval Using Bloom Filters

2016-07-12 · Araujo Andre, Chaves Jason, Lakshman Haricharan, Angst Roland 외

We consider the problem of using image queries to retrieve videos from a database. Our focus is on large-scale applications, where it is infeasible to index each database video frame independently. Our main contribution …

RetrievalVideo Retrieval

Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

2026-05-04 · Yiming Ding, Siyu Cao, Luyuan Jiao, Yixuan Li 외 arxiv

Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each query. This assumption does not always h…

Moment Retrieval

Mitigating Semantic Collapse in Partially Relevant Video Retrieval

2025-10-31 · WonJun Moon, MinSeok Jung, Gilhan Park, Tae-Young Kim 외 arxiv

Partially Relevant Video Retrieval (PRVR) seeks videos where only part of the content matches a text query. Existing methods treat every annotated text-video pair as a positive and all others as negatives, ignoring the r…

Partially Relevant Video RetrievalVideo Alignment

Unsupervised Segmentation of Action Segments in Egocentric Videos using Gaze

2017-09-30 · I. Hipiny, H. Ujir, J. L. Minoi, S. F. Samson Juan 외

Unsupervised segmentation of action segments in egocentric videos is a desirable feature in tasks such as activity recognition and content-based video retrieval. Reducing the search space into a finite set of action segm…

Activity RecognitionRetrievalVideo Retrieval

Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning

2025-09-01 · Long Zhang, Peipei Song, Jianfeng Dong, Kun Li 외 arxiv

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos partially relevant to a given query. The core challenge lies in learning robust query-video alignment against spurious semantic correlations ari…

Partially Relevant Video RetrievalVideo Alignment