paper-with-me

홈 › Papers

PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling

2026-06-24 · Zhijie Zheng, Xinhao Xiang, Jiawei Zhang arxiv

Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and content creation. Recent methods achieve outstanding performance by leveraging camera-controlled video diffusion models, but rely on iterative diffusion sampling, which greatly limits their practical deployment. We observe that geometric forward warping alone can cover the majority of a target view directly from the input image, with only a compact residual left for the encoder to correct. Motivated by this observation, we propose PRISM, a feed-forward framework that decomposes multi-view latent prediction into a parameter-free geometric prior and a learned residual correction, with no diffusion sampling required at inference. To enable generalization from purely synthetic training data, we devise a two-stage training strategy combining latents supervised distillation for geometric generalization and perceptual fine-tuning for appearance quality optimization. Extensive experiments on three benchmarks demonstrate that PRISM achieves competitive reconstruction quality compared with diffusion-based methods, while reducing inference time dramatically to only 36 seconds per scene.

📄 PDF Abstract BibTeX arXiv:2606.25430

Code (0)

등록된 구현이 없습니다.

Tasks

3D Reconstruction

Similar Papers 제목 키워드 기반

FHDR: HDR Image Reconstruction from a Single LDR Image using Feedback Network

2019-12-24 · Zeeshan Khan, Mukul Khanna, Shanmuganathan Raman

High dynamic range (HDR) image generation from a single exposure low dynamic range (LDR) image has been made possible due to the recent advances in Deep Learning. Various feed-forward Convolutional Neural Networks (CNNs)…

Image GenerationImage ReconstructionSingle-Image-Based Hdr Reconstruction

PRISM: Progressive Reasoning through Iterative Slot Memory for Vision

2026-05-29 · Ziyu Wang, Shuangpeng Han, Mengmi Zhang arxiv

Modern vision models process images in a single feed-forward pass, which limits their ability to recover missing evidence or refine uncertain representations under incomplete observations. Inspired by the iterative natur…

Semantic SegmentationImage ClassificationObject Detection

MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar Reconstruction

2026-04-23 · Yisheng He, Steven Hoi arxiv

We introduce MeshLAM, a feed-forward framework for one-shot animatable mesh head reconstruction that generates high-fidelity, animatable 3D head avatars from a single image. Unlike previous work that relies on time-consu…

Computational Efficiency

Perfusion Imaging and Single Material Reconstruction in Polychromatic Photon Counting CT

2026-02-02 · Namhoon Kim, Ashwin Pananjady, Amir Pourmorteza, Sara Fridovich-Keil arxiv

Background: Perfusion computed tomography (CT) images the dynamics of a contrast agent through the body over time, and is one of the highest X-ray dose scans in medical imaging. Recently, a theoretically justified recons…

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

2026-03-04 · Haian Jin, Rundi Wu, Tianyuan Zhang, Ruiqi Gao 외 arxiv

Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and $π^3$ have a computational cost that scales quadratically with the number of input images, making the…

3D Reconstruction