paper-with-me

Papers

LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction

2025-12-15 · Tianye Ding, Yiming Xie, Yiqing Liang, Moitreya Chatterjee, Pedro Miraldo, Huaizu Jiang arxiv

Recent feed-forward reconstruction models like VGGT and $π^3$ achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, limiting their practical deployment. While existing streaming methods address this through learned memory mechanisms or causal attention, they require extensive retraining and may not fully leverage the strong geometric priors of state-of-the-art offline models. We propose LASER, a training-free framework that converts an offline reconstruction model into a streaming system by aligning predictions across consecutive temporal windows. We observe that simple similarity transformation ($\mathrm{Sim}(3)$) alignment fails due to layer depth misalignment: monocular scale ambiguity causes relative depth scales of different scene layers to vary inconsistently between windows. To address this, we introduce layer-wise scale alignment, which segments depth predictions into discrete layers, computes per-layer scale factors, and propagates them across both adjacent windows and timestamps. Extensive experiments show that LASER achieves state-of-the-art performance on camera pose estimation and point map reconstruction %quality with offline models while operating at 14 FPS with 6 GB peak memory on a RTX A6000 GPU, enabling practical deployment for kilometer-scale streaming videos. Project website: https://neu-vi.github.io/LASER/

📄 PDF Abstract BibTeX arXiv:2512.13680

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose Estimation

Similar Papers 제목 키워드 기반

Very Low Resource Sentence Alignment: Luhya and Swahili

2022-10-01 · loresmt (COLING) 2022 10 · Everlyn Chimoto, Bruce Bassett

Language-agnostic sentence embeddings generated by pre-trained models such as LASER and LaBSE are attractive options for mining large datasets to produce parallel corpora for low-resource machine translation. We test LAS…

Machine TranslationSentenceSentence EmbeddingSentence-Embedding+1

Very Low Resource Sentence Alignment: Luhya and Swahili

2022-10-31 · Everlyn Asiko Chimoto, Bruce A. Bassett

Language-agnostic sentence embeddings generated by pre-trained models such as LASER and LaBSE are attractive options for mining large datasets to produce parallel corpora for low-resource machine translation. We test LAS…

Machine TranslationSentenceSentence EmbeddingSentence-Embedding+1

Shortest Trajectory of a Dubins Vehicle with a Controllable Laser

2024-03-19 · Shivam Bajaj, Bhargav Jha, Shaunak D. Bopardikar, Alexander Von Moll 외

We formulate a novel planar motion planning problem for a Dubins-Laser system that consists of a Dubins vehicle with an attached controllable laser. The vehicle moves with unit speed and the laser, having a finite range,…

Motion PlanningPosition

Automatic Scale Estimation of Structure from Motion based 3D Models using Laser Scalers

2019-06-19 · Klemen Istenic, Nuno Gracias, Aurelien Arnaubec, Javier Escartin 외

Recent advances in structure-from-motion techniques are enabling many scientific fields to benefit from the routine creation of detailed 3D models. However, for a large number of applications, only a single camera is ava…

Beyond Static Cropping: Layer-Adaptive Visual Localization and Decoding Enhancement

2026-02-04 · Zipeng Zhu, Zhanghao Hu, Qinglin Zhu, Yuxi Hong 외 arxiv

Large Vision-Language Models (LVLMs) have advanced rapidly by aligning visual patches with the text embedding space, but a fixed visual-token budget forces images to be resized to a uniform pretraining resolution, often …

Visual LocalizationQuestion AnsweringObject RecognitionVisual Grounding