paper-with-me

홈 › Papers

Seeing Across Time and Views: Multi-Temporal Cross-View Learning for Robust Video Person Re-Identification

2025-11-04 · Md Rashidunnabi, Kailash A. Hambarde, Vasco Lopes, Joao C. Neves, Hugo Proenca arxiv

Video-based person re-identification (ReID) in cross-view domains (for example, aerial-ground surveillance) remains an open problem because of extreme viewpoint shifts, scale disparities, and temporal inconsistencies. To address these challenges, we propose MTF-CVReID, a parameter-efficient framework that introduces seven complementary modules over a ViT-B/16 backbone. Specifically, we include: (1) Cross-Stream Feature Normalization (CSFN) to correct camera and view biases; (2) Multi-Resolution Feature Harmonization (MRFH) for scale stabilization across altitudes; (3) Identity-Aware Memory Module (IAMM) to reinforce persistent identity traits; (4) Temporal Dynamics Modeling (TDM) for motion-aware short-term temporal encoding; (5) Inter-View Feature Alignment (IVFA) for perspective-invariant representation alignment; (6) Hierarchical Temporal Pattern Learning (HTPL) to capture multi-scale temporal regularities; and (7) Multi-View Identity Consistency Learning (MVICL) that enforces cross-view identity coherence using a contrastive learning paradigm. Despite adding only about 2 million parameters and 0.7 GFLOPs over the baseline, MTF-CVReID maintains real-time efficiency (189 FPS) and achieves state-of-the-art performance on the AG-VPReID benchmark across all altitude levels, with strong cross-dataset generalization to G2A-VReID and MARS datasets. These results show that carefully designed adapter-based modules can substantially enhance cross-view robustness and temporal consistency without compromising computational efficiency. The source code is available at https://github.com/MdRashidunnabi/MTF-CVReID

📄 PDF Abstract BibTeX arXiv:2511.02564

Code (0)

등록된 구현이 없습니다.

Tasks

Person Re-IdentificationComputational EfficiencyContrastive Learning

Similar Papers 제목 키워드 기반

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

2024-12-04 · Hannan Lu, Xiaohe Wu, Shudong Wang, Xiameng Qin 외

Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Existing methods typically apply decoupled at…

Autonomous DrivingVideo Generation

InfiniteNature-Zero: Learning Perpetual View Generation of Natural Scenes from Single Images

2022-07-22 · Zhengqi Li, Qianqian Wang, Noah Snavely, Angjoo Kanazawa

We present a method for learning to generate unbounded flythrough videos of natural scenes starting from a single view, where this capability is learned from a collection of single photographs, without requiring camera p…

Perpetual View Generation

CROVIA: Seeing Drone Scenes from Car Perspective via Cross-View Adaptation

2023-04-14 · Thanh-Dat Truong, Chi Nhan Duong, Ashley Dowling, Son Lam Phung 외

Understanding semantic scene segmentation of urban scenes captured from the Unmanned Aerial Vehicles (UAV) perspective plays a vital role in building a perception model for UAV. With the limitations of large-scale densel…

Autonomous DrivingScene SegmentationSegmentation

TimeNeRF: Building Generalizable Neural Radiance Fields across Time from Few-Shot Input Views

2025-07-18 · Hsiang-Hui Hung, Huu-Phu Do, Yung-Hui Li, Ching-Chun Huang arxiv

We present TimeNeRF, a generalizable neural rendering approach for rendering novel views at arbitrary viewpoints and at arbitrary times, even with few input views. For real-world applications, it is expensive to collect …

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

2026-08-27 · Haizhao Fan, Xinyi Le arxiv

Time series forecasting models operate on raw numerical sequences, lacking the semantic knowledge that domain experts implicitly leverage, such as the physical meaning of each variable, its statistical behavior, and its …

Time Series Forecasting