HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video
Uncalibrated volumetric video streaming for human reconstruction is essential for holographic communication and AR/VR, yet remains challenging due to the need for temporal consistency and computational efficiency from sparse-view inputs. Existing methods rely on per-scene optimization or calibrated cameras, while recent feed-forward models are limited to low-resolution (0.5K) single-frame synthesis. We present HiReFF, a feed-forward method for 2K-resolution 360° human video reconstruction from uncalibrated sparse-view videos. Our framework decomposes the problem into two key tasks: foreground 3D Gaussian reconstruction from sparse-view videos (four views separated by 90°) and computationally efficient high-resolution synthesis. To enable the former, we propose Scale-synchronized Camera Calibration to resolve scale ambiguity for multi-view supervision, and Gaussian-wise Foreground Masking to reconstruct clean foregrounds by modulating Gaussian parameters. For efficient high-resolution synthesis, our High-resolution Side-tuning achieves 2K rendering by augmenting the Gaussian head with supplementary features while keeping the backbone at 0.5K, drastically reducing computational overhead. Experiments demonstrate that HiReFF significantly outperforms existing methods in high-resolution streaming volumetric video reconstruction. https://iridescentjiang.github.io/HiReFF
Code (0)
등록된 구현이 없습니다.
Tasks
Computational EfficiencyVideo ReconstructionSimilar Papers 제목 키워드 기반
SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors
Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches leverage generalizable models to generate 3…
3D ReconstructionPoint CloudsVideo Super-Resolution via Deep Draft-Ensemble Learning
We propose a new direction for fast video super-resolution (VideoSR) via a SR draft ensemble, which is defined as the set of high-resolution patch candidates before final image deconvolution. Our method contains two main…
Ensemble LearningImage DeconvolutionSuper-ResolutionVideo Super-ResolutionL2D2-GS: Learning to Densify for Feedforward Dynamic Gaussian Scene Reconstruction
High-fidelity reconstruction of dynamic urban environments is a cornerstone of autonomous driving simulation and large-scale world modeling. While 3D Gaussian Splatting (3DGS) has established a new standard for real-time…
Zero-shot GeneralizationAutonomous DrivingSCube: Instant Large-Scale Scene Reconstruction using VoxSplats
We present SCube, a novel method for reconstructing large-scale 3D scenes (geometry, appearance, and semantics) from a sparse set of posed images. Our method encodes reconstructed scenes using a novel representation VoxS…
3D ReconstructionScene GenerationPointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their per-pixel Gaussian prediction paradigm oft…
Autonomous Driving