paper-with-me

Papers

LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

2026-03-03 · Junyi Zhang, Charles Herrmann, Junhwa Hur, Chen Sun, Ming-Hsuan Yang, Forrester Cole, Trevor Darrell, Deqing Sun arxiv

Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadratic attention complexity or limited effective memory in recurrent designs. We present LoGeR (Long-context Geometric Reconstruction), a novel architecture that scales dense 3D reconstruction to extremely long sequences without post-optimization. LoGeR processes video streams in chunks, leveraging strong bidirectional priors for high-fidelity intra-chunk reasoning. To manage the critical challenge of coherence across chunk boundaries, we propose a learning-based hybrid memory module. This dual-component system combines a parametric Test-Time Training (TTT) memory to anchor the global coordinate frame and prevent scale drift, alongside a non-parametric Sliding Window Attention (SWA) mechanism to preserve uncompressed context for high-precision adjacent alignment. Remarkably, this memory architecture enables LoGeR to be trained on sequences of 128 frames, and generalize up to thousands of frames during inference. Evaluated across standard benchmarks and a newly repurposed VBR dataset with sequences of up to 19k frames, LoGeR substantially outperforms prior state-of-the-art feedforward methods--reducing ATE on KITTI by over 74%--and achieves robust, globally consistent reconstruction over unprecedented horizons.

📄 PDF Abstract BibTeX arXiv:2603.03269

Code (0)

등록된 구현이 없습니다.

Tasks

3D Reconstruction

Similar Papers 제목 키워드 기반

Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs

2026-07-01 · Chahan Vidal-Gorène, Nadi Tomeh, Victoria Khurshudyan arxiv

This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We introduce a new annotated dataset of 66 pages and compare geometric h…

Mem3R: Streaming 3D Reconstruction with Hybrid Memory via Test-Time Training

2026-04-08 · Changkun Liu, Jiezhi Yang, Zeman Li, Yuan Deng 외 arxiv

Streaming 3D perception is well suited to robotics and augmented reality, where long visual streams must be processed efficiently and consistently. Recent recurrent models offer a promising solution by maintaining fixed-…

3D ReconstructionDepth Estimation

RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation

2026-09-02 · Xiaolei Lang, Ze Kang, Zehao Huang, Naiyan Wang arxiv

Novel view synthesis from sparse inputs requires both geometric grounding from the observed views and generative priors of unobserved regions, motivating recent hybrid methods that combine reconstruction and generation. …

Novel View Synthesis

Geometric Context Transformer for Streaming 3D Reconstruction

2026-04-15 · Lin-Zhuo Chen, Jian Gao, Yihang Chen, Ka Leong Cheng 외 arxiv

Streaming 3D reconstruction aims to recover 3D information, such as camera poses and point clouds, from a video stream, which necessitates geometric accuracy, temporal consistency, and computational efficiency. Motivated…

Computational Efficiency3D ReconstructionPoint Clouds

Hybrid Channel Modeling and Environment Reconstruction for Terahertz Monostatic Sensing

2024-11-12 · Yejian Lyu, Zeyu Huang, Stefan Schwarz, Chong Han

THz ISAC aims to integrate novel functionalities, such as positioning and environmental sensing, into communication systems. Accurate channel modeling is crucial for the design and performance evaluation of future ISAC s…

ISAC