paper-with-me

홈 › Papers

PointSt3R: Point Tracking through 3D Grounded Correspondence

2025-10-30 · Rhodri Guerrier, Adam W. Harley, Dima Damen arxiv

Recent advances in foundational 3D reconstruction models, such as DUSt3R and MASt3R, have shown great potential in 2D and 3D correspondence in static scenes. In this paper, we propose to adapt them for the task of point tracking through 3D grounded correspondence. We first demonstrate that these models are competitive point trackers when focusing on static points, present in current point tracking benchmarks ($+33.5\%$ on EgoPoints vs. CoTracker2). We propose to combine the reconstruction loss with training for dynamic correspondence along with a visibility head, and fine-tuning MASt3R for point tracking using a relatively small amount of synthetic data. Importantly, we only train and evaluate on pairs of frames where one contains the query point, effectively removing any temporal context. Using a mix of dynamic and static point correspondences, we achieve competitive or superior point tracking results on four datasets (e.g. competitive on TAP-Vid-DAVIS 73.8 $δ_{avg}$ / 85.8\% occlusion acc. for PointSt3R compared to 75.7 / 88.3\% for CoTracker2; and significantly outperform CoTracker3 on EgoPoints 61.3 vs 54.2 and RGB-S 87.0 vs 82.8). We also present results on 3D point tracking along with several ablations on training datasets and percentage of dynamic correspondences.

📄 PDF Abstract BibTeX arXiv:2510.26443

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionPoint Tracking

Results from the Paper

RankTaskDatasetModelMetrics
#7 Point Tracking TAP-Vid-DAVIS PointSt3R Occlusion Accuracy: 88.3

Similar Papers 제목 키워드 기반

Advanced Feature Learning on Point Clouds using Multi-resolution Features and Learnable Pooling

2022-05-20 · Kevin Tirta Wijaya, Dong-Hee Paek, Seung-Hyun Kong

Existing point cloud feature learning networks often incorporate sequences of sampling, neighborhood grouping, neighborhood-wise feature learning, and feature aggregation to learn high-semantic point features that repres…

3D Point Cloud Classification

SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models

2026-05-29 · Olaf Dünkel, Basavaraj Sunagad, Haoran Wang, David T. Hoffmann 외 arxiv

Measuring structured object understanding in vision foundation models remains challenging due to inconsistent evaluation protocols and limited part-level supervision. Semantic correspondence (SC) evaluates this capabilit…

Semantic correspondence3D Pose EstimationImage Matching

3AM: 3egment Anything with Geometric Consistency in Videos

2026-01-13 · Yang-Che Sun, Cheng Sun, Chin-Yang Lin, Fu-En Yang 외 arxiv

Video object segmentation methods like SAM2 achieve strong performance through memory-based architectures but struggle under large viewpoint changes due to reliance on appearance features. Traditional 3D instance segment…

Video Object Segmentation3D Instance Segmentation

Autogenic Language Embedding for Coherent Point Tracking

2024-07-30 · Zikai Song, Ying Tang, Run Luo, Lintao Ma 외

Point tracking is a challenging task in computer vision, aiming to establish point-wise correspondence across long video sequences. Recent advancements have primarily focused on temporal modeling techniques to improve lo…

DecoderPoint Tracking

Self-supervised Keypoint Correspondences for Multi-Person Pose Estimation and Tracking in Videos

2020-04-27 · ECCV 2020 8 · Umer Rafi, Andreas Doering, Bastian Leibe, Juergen Gall

Video annotation is expensive and time consuming. Consequently, datasets for multi-person pose estimation and tracking are less diverse and have more sparse annotations compared to large scale image datasets for human po…

Multi-Person Pose EstimationMulti-Person Pose Estimation and TrackingPose EstimationPose Tracking