paper-with-me

홈 › Papers

Circulant temporal encoding for video retrieval and temporal alignment

2015-06-08 · Matthijs Douze, Jérôme Revaud, Jakob Verbeek, Hervé Jégou, Cordelia Schmid

We address the problem of specific video event retrieval. Given a query video of a specific event, e.g., a concert of Madonna, the goal is to retrieve other videos of the same event that temporally overlap with the query. Our approach encodes the frame descriptors of a video to jointly represent their appearance and temporal order. It exploits the properties of circulant matrices to efficiently compare the videos in the frequency domain. This offers a significant gain in complexity and accurately localizes the matching parts of videos. The descriptors can be compressed in the frequency domain with a product quantizer adapted to complex numbers. In this case, video retrieval is performed without decompressing the descriptors. We also consider the temporal alignment of a set of videos. We exploit the matching confidence and an estimate of the temporal offset computed for all pairs of videos by our retrieval approach. Our robust algorithm aligns the videos on a global timeline by maximizing the set of temporally consistent matches. The global temporal alignment enables synchronous playback of the videos of a given scene.

📄 PDF Abstract BibTeX arXiv:1506.02588

Code (1)

facebookresearch/videoalignment pytorch

Tasks

RetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

Event Retrieval in Large Video Collections with Circulant Temporal Encoding

2013-06-01 · CVPR 2013 6 · Jerome Revaud, Matthijs Douze, Cordelia Schmid, Herve Jegou

This paper presents an approach for large-scale event retrieval. Given a video clip of a specific event, e.g., the wedding of Prince William and Kate Middleton, the goal is to retrieve other videos representing the same …

Copy DetectionQuantizationRetrieval

3D Pose from Motion for Cross-view Action Recognition via Non-linear Circulant Temporal Encoding

2014-06-01 · CVPR 2014 6 · Ankur Gupta, Julieta Martinez, James J. Little, Robert J. Woodham

We describe a new approach to transfer knowledge across views for action recognition by using examples from a large collection of unlabelled mocap data. We achieve this by directly matching purely motion based features f…

Action RecognitionTemporal Action Localization

Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval

2020-07-06 · Xun Yang, Jianfeng Dong, Yixin Cao, Xun Wang 외

The rapid growth of user-generated videos on the Internet has intensified the need for text-based video retrieval systems. Traditional methods mainly favor the concept-based paradigm on retrieval with simple queries, whi…

RetrievalVideo Retrieval

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

2025-06-17 · Xiaoqi Wang, Yi Wang, Lap-Pui Chau

Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-trai…

Multi-Instance RetrievalRetrievalVideo Understanding

TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

2023-10-29 · Shuhuai Ren, Sishuo Chen, Shicheng Li, Xu sun 외

Large-scale video-language pre-training has made remarkable strides in advancing video-language understanding tasks. However, the heavy computational burden of video encoding remains a formidable efficiency bottleneck, p…

FormLanguage ModellingRetrievalVideo Question Answering+2