paper-with-me

홈 › Papers

RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D Reconstruction

2025-07-23 · Yuqing Lan, Chenyang Zhu, Shuaifeng Zhi, Jiazhao Zhang, Zhoufeng Wang, Renjiao Yi, Yijie Wang, Kai Xu arxiv

The introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it improves the mapping completeness and memory efficiency. However, the lack of reconstruction details and the time-consuming learning of neural representations hinder the widespread application of neural-based methods to large-scale online reconstruction. We introduce RemixFusion, a novel residual-based mixed representation for scene reconstruction and camera pose estimation dedicated to high-quality and large-scale online RGB-D reconstruction. In particular, we propose a residual-based map representation comprised of an explicit coarse TSDF grid and an implicit neural module that produces residuals representing fine-grained details to be added to the coarse grid. Such mixed representation allows for detail-rich reconstruction with bounded time and memory budget, contrasting with the overly-smoothed results by the purely implicit representations, thus paving the way for high-quality camera tracking. Furthermore, we extend the residual-based representation to handle multi-frame joint pose optimization via bundle adjustment (BA). In contrast to the existing methods, which optimize poses directly, we opt to optimize pose changes. Combined with a novel technique for adaptive gradient amplification, our method attains better optimization convergence and global optimality. Furthermore, we adopt a local moving volume to factorize the mixed scene representation with a divide-and-conquer design to facilitate efficient online learning in our residual-based framework. Extensive experiments demonstrate that our method surpasses all state-of-the-art ones, including those based either on explicit or implicit representations, in terms of the accuracy of both mapping and tracking on large-scale scenes.

📄 PDF Abstract BibTeX arXiv:2507.17594

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose Estimation

Similar Papers 제목 키워드 기반

CWRNN-INVR: A Coupled WarpRNN based Implicit Neural Video Representation

2026-04-08 · Yiyang Li, Yanbo Gao, Shuai Li, Zhenyu Du 외 arxiv

Implicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures ef…

Triple Attention Mixed Link Network for Single Image Super Resolution

2018-10-08 · Xi Cheng, Xiang Li, Jian Yang

Single image super resolution is of great importance as a low-level computer vision task. Recent approaches with deep convolutional neural networks have achieved im-pressive performance. However, existing architectures h…

Image Super-ResolutionSuper-Resolution

Interpreting the Residual Stream of ResNet18

2024-07-07 · André Longon

A mechanistic understanding of the computations learned by deep neural networks (DNNs) is far from complete. In the domain of visual object recognition, prior research has illuminated inner workings of InceptionV1, but D…

Object Recognition

Towards Understanding Residual and Dilated Dense Neural Networks via Convolutional Sparse Coding

2019-12-05 · Zhiyang Zhang, Shihua Zhang

Convolutional neural network (CNN) and its variants have led to many state-of-art results in various fields. However, a clear theoretical understanding about them is still lacking. Recently, multi-layer convolutional spa…

Naturally Computed Scale Invariance in the Residual Stream of ResNet18

2025-04-22 · André Longon

An important capacity in visual object recognition is invariance to image-altering variables which leave the identity of objects unchanged, such as lighting, rotation, and scale. How do neural networks achieve this? Prio…

Object Recognition