paper-with-me

Papers

Temporal Complementarity-Guided Reinforcement Learning for Image-to-Video Person Re-Identification

2022-01-01 · CVPR 2022 1 · Wei Wu, Jiawei Liu, Kecheng Zheng, Qibin Sun, Zheng-Jun Zha

Image-to-video person re-identification aims to retrieve the same pedestrian as the image-based query from a video-based gallery set. Existing methods treat it as a cross-modality retrieval task and learn the common latent embeddings from image and video modalities, which are both less effective and efficient due to large modality gap and redundant feature learning by utilizing all video frames. In this work, we first regard this task as point-to-set matching problem identical to human decision process, and propose a novel Temporal Complementarity-Guided Reinforcement Learning (TCRL) approach for image-to-video person re-identification. TCRL employs deep reinforcement learning to make sequential judgments on dynamically selecting suitable amount of frames from gallery videos, and accumulate adequate temporal complementary information among these frames by the guidance of the query image, towards balancing efficiency and accuracy. Specifically, TCRL formulates point-to-set matching procedure as Markov decision process, where a sequential judgement agent measures the uncertainty between the query image and all historical frames at each time step, and verifies that sufficient complementary clues are accumulated for judgment (same or different) or one more frames are requested to assist judgment. Moreover, TCRL maintains a sequential feature extraction module with a complementary residual detector to dynamically suppress redundant salient regions and thoroughly mine diverse complementary clues among these selected frames for enhancing frame-level representation. Extensive experiments demonstrate the superiority of our method.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningImage-To-Video Person Re-IdentificationPerson Re-Identificationreinforcement-learningReinforcement Learning (RL)Retrievalset matchingVideo-Based Person Re-Identification

Similar Papers 제목 키워드 기반

VideoFusion: A Spatio-Temporal Collaborative Network for Mutli-modal Video Fusion and Restoration

2025-03-30 · Linfeng Tang, Yeda Wang, Meiqi Gong, Zizhuo Li 외

Compared to images, videos better align with real-world acquisition scenarios and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multipl…

Sensor Fusion

TimeZero: Temporal Video Grounding with Reasoning-Guided LVLM

2025-03-17 · Ye Wang, Boshen Xu, Zihao Yue, Zihan Xiao 외

We introduce TimeZero, a reasoning-guided LVLM designed for the temporal video grounding (TVG) task. This task requires precisely localizing relevant video segments within long videos based on a given language query. Tim…

Video Grounding

Visible and NIR Image Fusion Algorithm Based on Information Complementarity

2023-09-19 · Zhuo Li, Bo Li

Visible and near-infrared(NIR) band sensors provide images that capture complementary spectral radiations from a scene. And the fusion of the visible and NIR image aims at utilizing their spectrum properties to enhance i…

Two-stream Collaborative Learning with Spatial-Temporal Attention for Video Classification

2017-11-09 · Yuxin Peng, Yunzhen Zhao, Junchao Zhang

Video classification is highly important with wide applications, such as video search and intelligent surveillance. Video naturally consists of static and motion information, which can be represented by frame and optical…

General ClassificationOptical Flow EstimationVideo ClassificationVocal Bursts Valence Prediction

Edit as You See: Image-guided Video Editing via Masked Motion Modeling

2025-01-08 · Zhi-Lin Huang, Yixuan Liu, Chujun Qin, Zhongdao Wang 외

Recent advancements in diffusion models have significantly facilitated text-guided video editing. However, there is a relative scarcity of research on image-guided video editing, a method that empowers users to edit vide…

Optical Flow EstimationSelf-Supervised LearningVideo Editing