paper-with-me

홈 › Papers

Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels

2026-03-03 · Jiahao Lu, Jiayi Xu, Wenbo Hu, Ruijie Zhu, Chengfeng Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu arxiv

Estimating the 3D trajectory of every pixel from a monocular video is crucial and promising for a comprehensive understanding of the 3D dynamics of videos. Recent monocular 3D tracking works demonstrate impressive performance, but are limited to either tracking sparse points on the first frame or a slow optimization-based framework for dense tracking. In this paper, we propose a feedforward model, called Track4World, enabling an efficient holistic 3D tracking of every pixel in the world-centric coordinate system. Built on the global 3D scene representation encoded by a VGGT-style ViT, Track4World applies a novel 3D correlation scheme to simultaneously estimate the pixel-wise 2D and 3D dense flow between arbitrary frame pairs. The estimated scene flow, along with the reconstructed 3D geometry, enables subsequent efficient 3D tracking of every pixel of this video. Extensive experiments on multiple benchmarks demonstrate that our approach consistently outperforms existing methods in 2D/3D flow estimation and 3D tracking, highlighting its robustness and scalability for real-world 4D reconstruction tasks.

📄 PDF Abstract BibTeX arXiv:2603.02573

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels

2025-12-09 · Jiahao Lu, Weitao Xiong, Jiacheng Deng, Peng Li 외 arxiv

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking m…

Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow

2026-02-15 · Shenhan Qian, Ganlin Zhang, Shangzhe Wu, Daniel Cremers arxiv

Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motion: static multi-view reconstruction systems assume a rigid world, wherea…

Camera Pose EstimationScene Understanding

Instance Tracking in 3D Scenes from Egocentric Videos

2023-12-07 · CVPR 2024 1 · Yunhan Zhao, Haoyu Ma, Shu Kong, Charless Fowlkes

Egocentric sensors such as AR/VR devices capture human-object interactions and offer the potential to provide task-assistance by recalling 3D locations of objects of interest in the surrounding environment. This capabili…

Human-Object Interaction DetectionObject Tracking

OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and Captioning

2024-04-04 · Anwesa Choudhuri, Girish Chowdhary, Alexander G. Schwing

We propose the new task 'open-world video instance segmentation and captioning'. It requires to detect, segment, track and describe with rich captions never before seen objects. This challenging task can be addressed by …

DescriptiveDiversityInstance SegmentationLanguage Modeling+7

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

2025-05-08 · Henry Zheng, Hao Shi, Qihang Peng, Yong Xien Chng 외

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D vi…

3D visual groundingcross-modal alignmentVisual Grounding