paper-with-me

Papers

PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization

2025-02-11 · Bing Fan, Yunhe Feng, Yapeng Tian, Yuewei Lin, Yan Huang, Heng Fan

Egocentric visual query localization (EgoVQL) focuses on localizing the target of interest in space and time from first-person videos, given a visual query. Despite recent progressive, existing methods often struggle to handle severe object appearance changes and cluttering background in the video due to lacking sufficient target cues, leading to degradation. Addressing this, we introduce PRVQL, a novel Progressive knowledge-guided Refinement framework for EgoVQL. The core is to continuously exploit target-relevant knowledge directly from videos and utilize it as guidance to refine both query and video features for improving target localization. Our PRVQL contains multiple processing stages. The target knowledge from one stage, comprising appearance and spatial knowledge extracted via two specially designed knowledge learning modules, are utilized as guidance to refine the query and videos features for the next stage, which are used to generate more accurate knowledge for further feature refinement. With such a progressive process, target knowledge in PRVQL can be gradually improved, which, in turn, leads to better refined query and video features for localization in the final stage. Compared to previous methods, our PRVQL, besides the given object cues, enjoys additional crucial target information from a video as guidance to refine features, and hence enhances EgoVQL in complicated scenes. In our experiments on challenging Ego4D, PRVQL achieves state-of-the-art result and largely surpasses other methods, showing its efficacy. Our code, model and results will be released at https://github.com/fb-reps/PRVQL.

📄 PDF Abstract BibTeX arXiv:2502.07707

Code (1)

fb-reps/prvql 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

2025-03-12 · Haoyu Zhang, Qiaohui Chu, Meng Liu, Yunxiao Wang 외

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. Current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exoce…

Instruction FollowingVideo Understanding

EgoEV-HandPose: Egocentric 3D Hand Pose Estimation and Gesture Recognition with Stereo Event Cameras

2026-05-12 · Luming Wang, Hao Shi, Jiajun Zhai, Kailun Yang 외 arxiv

Egocentric 3D hand pose estimation and gesture recognition are essential for immersive augmented/virtual reality, human-computer interaction, and robotics. However, conventional frame-based cameras suffer from motion blu…

3D Hand Pose EstimationGesture Recognition

TSR-Ego: Temporally Guided Stereo Refinement Framework for Egocentric 3D Human Pose Estimation

2026-07-10 · Md Mushfiqur Azam, John Quarles, Kevin Desai arxiv

Egocentric 3D human pose estimation from head-mounted stereo cameras is challenging due to fisheye distortion, severe self-occlusion, and frequent truncation of body joints outside the camera field of view. Recent stereo…

3D Human Pose EstimationPose Prediction

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting

2026-07-17 · Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Xiang Li 외 arxiv

Perceiving multimodal cues and forecasting fine-grained actions from an egocentric (Ego) perspective is vital for applications like robot manipulation. However, previous studies either rely mainly on under-informed visua…

Robot Manipulation

Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse

2026-09-09 · Heng Zhu, Ye Liu, Kun Zhu, Feifei Song arxiv

Existing knowledge-base-assisted semantic communication schemes commonly adopt either single knowledge-base quantization (SKBQ) or multi-knowledge-base residual quantization (MKBQ). SKBQ incurs limited storage overhead b…

Semantic Communication