3D-LMVIC: Learning-based Multi-View Image Coding with 3D Gaussian Geometric Priors
Existing multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camera systems, commonly found in virtual reality and autonomous driving applications. To address this limitation, we propose 3D-LMVIC, a novel learning-based multi-view image compression framework that leverages 3D Gaussian Splatting to derive geometric priors for accurate disparity estimation. Furthermore, we introduce a depth map compression model to minimize geometric redundancy across views, along with a multi-view sequence ordering strategy based on a defined distance measure between views to enhance correlations between adjacent views. Experimental results demonstrate that 3D-LMVIC achieves superior performance compared to both traditional and learning-based methods. Additionally, it significantly improves disparity estimation accuracy over existing two-view approaches.
Code (1)
Tasks
Autonomous DrivingDisparity EstimationImage CompressionSimilar Papers 제목 키워드 기반
LocusGS: Spatially Grounded Tokens for Feed-Forward 3D Gaussian Splatting
Recent query-based feed-forward 3DGS methods represent a scene using learnable queries, each aggregating multi-view evidence and decoding a group of Gaussians. Ideally, different queries should specialize in coherent loc…
Novel View SynthesisSingle-View 3D Reconstruction via SO(2)-Equivariant Gaussian Sculpting Networks
This paper introduces SO(2)-Equivariant Gaussian Sculpting Networks (GSNs) as an approach for SO(2)-Equivariant 3D object reconstruction from single-view image observations. GSNs take a single observation as input to gen…
3D Object Reconstruction3D ReconstructionObjectObject Reconstruction+1Turbo3D: Ultra-fast Text-to-3D Generation
We present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forwa…
3D GenerationText to 3DMulti-View Stereo by Temporal Nonparametric Fusion
We propose a novel idea for depth estimation from multi-view image-pose pairs, where the model has capability to leverage information from previous latent-space encodings of the scene. This model uses pairs of images and…
DecoderDepth EstimationDisparity EstimationGAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion
We propose a novel approach for reconstructing animatable 3D Gaussian avatars from monocular videos captured by commodity devices like smartphones. Photorealistic 3D head avatar reconstruction from such recordings is cha…
Novel View SynthesisSSIM