Connecting the Complementary-View Videos: Joint Camera Identification and Subject Association
We attempt to connect the data from complementary views, i.e., top view from drone-mounted cameras in the air, and side view from wearable cameras on the ground. Collaborative analysis of such complementary-view data can facilitate to build the air-ground cooperative visual system for various kinds of applications. This is a very challenging problem due to the large view difference between top and side views. In this paper, we develop a new approach that can simultaneously handle three tasks: i) localizing the side-view camera in the top view; ii) estimating the view direction of the side-view camera; iii) detecting and associating the same subjects on the ground across the complementary views. Our main idea is to explore the spatial position layout of the subjects in two views. In particular, we propose a spatial-aware position representation method to embed the spatial-position distribution of the subjects in different views. We further design a cross-view video collaboration framework composed of a camera identification module and a subject association module to simultaneously perform the above three tasks. We collect a new synthetic dataset consisting of top-view and side-view video sequence pairs for performance evaluation and the experimental results show the effectiveness of the proposed method.
Code (1)
Tasks
PositionSimilar Papers 제목 키워드 기반
CamWorldQA: Perceptual Quality Assessment of Camera-Controlled World Video Generation
Recent advances in generative video models have enabled camera-controlled world video generation, allowing models to synthesize videos under user-defined camera trajectories. However, existing video quality assessment (V…
Video Quality AssessmentVideo Generation3D Human Pose Estimation in Multi-View Operating Room Videos Using Differentiable Camera Projections
3D human pose estimation in multi-view operating room (OR) videos is a relevant asset for person tracking and action recognition. However, the surgical environment makes it challenging to find poses due to sterile clothi…
3D Human Pose EstimationAction RecognitionPose EstimationIntegrating Egocentric Videos in Top-view Surveillance Videos: Joint Identification and Temporal Alignment
Videos recorded from first person (egocentric) perspective have little visual appearance in common with those from third person perspective, especially with videos captured by top-view surveillance cameras. In this paper…
Joint Person Segmentation and Identification in Synchronized First- and Third-person Videos
In a world of pervasive cameras, public spaces are often captured from multiple perspectives by cameras of different types, both fixed and mobile. An important problem is to organize these heterogeneous collections of vi…
SegmentationSelf-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera
We present GLNet, a self-supervised framework for learning depth, optical flow, camera pose and intrinsic parameters from monocular video - addressing the difficulty of acquiring realistic ground-truth for such tasks. We…
Monocular Depth EstimationOptical Flow EstimationSelf-Supervised LearningTransfer Learning+1