paper-with-me

Papers

Visio-Temporal Attention for Multi-Camera Multi-Target Association

2021-01-01 · ICCV 2021 10 · Yu-Jhe Li, Xinshuo Weng, Yan Xu, Kris M. Kitani

We address the task of Re-Identification (Re-ID) in multi-target multi-camera (MTMC) tracking where we track multiple pedestrians using multiple overlapping uncalibrated (unknown pose) cameras. Since the videos are temporally synchronized and spatially overlapping, we can see a person from multiple views and associate their trajectory across cameras. In order to find the correct association between pedestrians visible from multiple views during the same time window, we extract a visual feature from a tracklet (sequence of pedestrian images) that encodes its similarity and dissimilarity to all other candidate tracklets. We propose a inter-tracklet (person to person) attention mechanism that learns a representation for a target tracklet while taking into account other tracklets across multiple views. Furthermore, to encode the gait and motion of a person, we introduce second intra-tracklet (person-specific) attention module with position embeddings. This second module employs a transformer encoder to learn a feature from a sequence of features over one tracklet. Experimental results on WILDTRACK and our new dataset `ConstructSite' confirm the superiority of our model over state-of-the-art ReID methods (5% and 10% performance gain respectively) in the context of uncalibrated MTMC tracking. While our model is designed for overlapping cameras, we also obtain state-of-the-art results on two other benchmark datasets (MARS and DukeMTMC) with non-overlapping cameras.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TRIG: Trajectory-Rig Decoupled Metric Geometry Learning

2026-07-07 · Lizhou Liao, Wentao Xu, Handong Wang, Lirong Yang 외 arxiv

Vision-centric autonomous driving requires accurate metric geometry and ego-motion estimation from synchronized multi-camera observations. Recent visual geometry models show strong performance in pose estimation, depth p…

Autonomous Driving3D ReconstructionPose Estimation

LMGP: Lifted Multicut Meets Geometry Projections for Multi-Camera Multi-Object Tracking

2021-11-23 · CVPR 2022 1 · Duy M. H. Nguyen, Roberto Henschel, Bodo Rosenhahn, Daniel Sonntag 외

Multi-Camera Multi-Object Tracking is currently drawing attention in the computer vision field due to its superior performance in real-world applications such as video surveillance in crowded scenes or in wide spaces. In…

3D geometryMulti-Object TrackingMultiple Object TrackingObject Tracking

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining

2026-07-05 · Jingyu Song, Yi Liu, Katherine A. Skinner arxiv

Camera-radar (CR) fusion is a practical sensing configuration for autonomous driving, but existing models are typically trained with task-specific supervision, limiting reusable representation learning. We present CRISP,…

Representation LearningAutonomous DrivingMotion ForecastingPoint Clouds

Doracamom: Joint 3D Detection and Occupancy Prediction with Multi-view 4D Radars and Cameras for Omnidirectional Perception

2025-01-26 · Lianqing Zheng, Jianan Liu, Runwei Guan, Long Yang 외

3D object detection and occupancy prediction are critical tasks in autonomous driving, attracting significant attention. Despite the potential of recent vision-based methods, they encounter challenges under adverse condi…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection+1

Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention

2024-10-14 · Dejia Xu, Yifan Jiang, Chen Huang, Liangchen Song 외

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to i…

Image to Video GenerationVideo Generation