paper-with-me

홈 › Papers

GeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric Assembly

2026-08-18 · Tinghao Jiang, Sheng Tang, Shengzhe Wei, Juntong Fang, Weiqi Zhang, Junsheng Zhou, Zesong Li arxiv

Long-sequence 3D reconstruction from RGB videos requires both accurate local geometry and globally consistent camera motion. Feed-forward models provide strong depth and pose predictions, but their memory cost prevents joint inference over long sequences. Chunk-wise processing improves scalability, yet independently predicted chunks often exhibit scale drift, pose errors, and point-cloud misalignment. We present GeoWeaver, a unified framework comprising a Geometric Prior Model (GPM) and Test-Time Adaptation (TTA). The GPM predicts chunk-wise depth, confidence, and camera parameters as adjustable geometric priors. TTA then performs sequential initialization, global chunk-level Sim(3) alignment, and coarse-to-fine refinement of camera poses, affine depth corrections, and intrinsics. Dense correspondences provide adjacent, cross-chunk, and long-range constraints, while a robust CDF-style objective jointly optimizes weighted 2D reprojection and 3D consistency residuals. This design preserves local geometric accuracy while correcting accumulated pose, scale, depth, and calibration errors. Experiments across diverse long-sequence benchmarks demonstrate improved camera accuracy, global consistency, and point-cloud quality. Ablations verify the contribution of each adaptation stage, and applying the same TTA procedure to different geometric prior models consistently improves their trajectory estimates, demonstrating that GeoWeaver is not tied to a specific GPM.

📄 PDF Abstract BibTeX arXiv:2608.17389

Code (0)

등록된 구현이 없습니다.

Tasks

Test-time Adaptation3D Reconstruction

Similar Papers 제목 키워드 기반

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

2026-05-21 · Deshui Miao, Xingsen Huang, Yameng Gu, Xin Li 외 arxiv

Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearance. Recent multimodal models incorporate geometric information throu…

Spatial Reasoning

LIST3R: Long-sequence Instance-aware 3D Reconstruction

2026-07-01 · Jing Gao, Wei Wang, Feiran Wang, Yan Yan arxiv

We present LIST3R, an instance-aware framework for long-sequence 3D reconstruction inspired by the way humans organize spatial memory around stable and recognizable objects. LIST3R organizes long-sequence reconstruction …

3D Reconstruction

GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D Reconstruction

2026-05-15 · Leyang Chen, Junyi Wu, Zhiteng Li, Yulun Zhang arxiv

Streaming 3D reconstruction from long monocular video sequences requires maintaining a key-value (KV) cache that grows linearly with sequence length, creating a severe memory bottleneck. Existing approaches either trunca…

3D Reconstruction

A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music

2018-03-13 · ICML 2018 7 · Adam Roberts, Jesse Engel, Colin Raffel, Curtis Hawthorne 외

The Variational Autoencoder (VAE) has proven to be an effective model for producing semantically meaningful latent representations for natural data. However, it has thus far seen limited application to sequential data, a…

Decoder

Joint Layout Estimation and Global Multi-View Registration for Indoor Reconstruction

2017-04-25 · ICCV 2017 10 · Jeong-Kyun Lee, Jae-Won Yea, Min-Gyu Park, Kuk-Jin Yoon

In this paper, we propose a novel method to jointly solve scene layout estimation and global registration problems for accurate indoor 3D reconstruction. Given a sequence of range data, we first build a set of scene frag…

3D ReconstructionClustering