paper-with-me

홈 › Papers

Any4D: Unified Feed-Forward Metric 4D Reconstruction

2025-12-11 · Jay Karhade, Nikhil Keetha, Yuchen Zhang, Tanisha Gupta, Akash Sharma, Sebastian Scherer, Deva Ramanan arxiv

We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N frames, in contrast to prior work that typically focuses on either 2-view dense scene flow or sparse 3D point tracking. Moreover, unlike other recent methods for 4D reconstruction from monocular RGB videos, Any4D can process additional modalities and sensors such as RGB-D frames, IMU-based egomotion, and Radar Doppler measurements, when available. One of the key innovations that allows for such a flexible framework is a modular representation of a 4D scene; specifically, per-view 4D predictions are encoded using a variety of egocentric factors (depthmaps and camera intrinsics) represented in local camera coordinates, and allocentric factors (camera extrinsics and scene flow) represented in global world coordinates. We achieve superior performance across diverse setups - both in terms of accuracy (2-3X lower error) and compute efficiency (15X faster), opening avenues for multiple downstream applications.

📄 PDF Abstract BibTeX arXiv:2512.10935

Code (0)

등록된 구현이 없습니다.

Tasks

Point Tracking

Similar Papers 제목 키워드 기반

UniQueR: Unified Query-based Feedforward 3D Reconstruction

2026-03-24 · Chensheng Peng, Quentin Herau, Jiezhi Yang, Yichen Xie 외 arxiv

We present UniQueR, a unified query-based feedforward framework for efficient and accurate 3D reconstruction from unposed images. Existing feedforward models such as DUSt3R, VGGT, and AnySplat typically predict per-pixel…

3D Reconstruction

MapAnything: Universal Feed-Forward Metric 3D Reconstruction

2025-09-16 · Nikhil Keetha, Norman Müller, Johannes Schönberger, Lorenzo Porzi 외 arxiv

We introduce MapAnything, a unified transformer-based feed-forward model that ingests one or more images along with optional geometric inputs such as camera intrinsics, poses, depth, or partial reconstructions, and then …

Monocular Depth EstimationCamera Localization3D ReconstructionDepth Completion

Pano3D: Unified 3D Reconstruction and Panoptic Segmentation

2026-06-12 · Victor Barberteguy, Ahmet Iscen, Mathilde Caron, Alireza Fathi 외 arxiv

Recent advances in 3D feedforward reconstruction neural networks have achieved remarkable success in dense reconstruction from images without any camera parameters. Yet, equipping these models with robust semantic unders…

Panoptic Segmentation3D Reconstruction

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

2025-11-28 · Tianling Xu, Shengzhe Gan, Leslie Gu, Yuelei Li 외 arxiv

Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images. However, exis…

3D Reconstruction

UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass

2026-01-03 · Mengfei Li, Peng Li, Zheng Zhang, Jiahao Lu 외 arxiv

We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance o…

Point Clouds