paper-with-me

Papers

Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus Cues

2025-01-01 · CVPR 2025 1 · Yuhui Liu, Liangxun Ou, Qiang Fu, Hadi Amata, Wolfgang Heidrich, Yifan Peng

Extracting high-fidelity RGBD information from two-dimensional (2D) images is essential for various visual computing applications. Stereo imaging, as a reliable passive imaging technique for obtaining three-dimensional (3D) scene information, has benefited greatly from deep learning advancements. However, existing stereo depth estimation algorithms struggle to perceive high-frequency information and resolve high-resolution depth maps in realistic camera settings with large depth variations. These algorithms commonly neglect the hardware parameter configuration, limiting the potential for achieving optimal solutions solely through software-based design strategies.This work presents a hardware-software co-designed RGBD imaging framework that leverages both stereo and focus cues to reconstruct texture-rich color images along with detailed depth maps over a wide depth range. A pair of rank-2 parameterized diffractive optical elements (DOEs) is employed to encode perpendicular complementary information optically during stereo acquisitions. Additionally, we employ an IGEV-UNet-fused neural network tailored to the proposed rank-2 encoding for stereo matching and image reconstruction. Through prototyping a stereo camera with customized DOEs, our deep stereo imaging paradigm has demonstrated superior performance over existing monocular and stereo imaging systems in both image PSNR by 2.96 dB gain and depth accuracy in high-frequency details across distances from 0.67 to 8 meters.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationImage ReconstructionStereo Depth EstimationStereo Matching

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Learned Off-aperture Encoding for Wide Field-of-view RGBD Imaging

2025-07-30 · Haoyu Wei, Xin Liu, Yuhui Liu, Qiang Fu 외 arxiv

End-to-end (E2E) designed imaging systems integrate coded optical designs with decoding algorithms to enhance imaging fidelity for diverse visual tasks. However, existing E2E designs encounter significant challenges in m…

Seeing Clearly and Deeply: An RGBD Imaging Approach with a Bio-inspired Monocentric Design

2025-10-29 · Zongxi Yu, Xiaolong Qian, Shaohua Gao, Qi Jiang 외 arxiv

Achieving high-fidelity, compact RGBD imaging presents a dual challenge: conventional compact optics struggle with RGB sharpness across the entire depth-of-field, while software-only Monocular Depth Estimation (MDE) is a…

Monocular Depth EstimationImage Restoration

AGG-Net: Attention Guided Gated-convolutional Network for Depth Image Completion

2023-09-04 · ICCV 2023 1 · Dongyue Chen, Tingxuan Huang, Zhimin Song, Shizhuo Deng 외

Recently, stereo vision based on lightweight RGBD cameras has been widely used in various fields. However, limited by the imaging principles, the commonly used RGB-D cameras based on TOF, structured light, or binocular v…

Stereo World Model: Camera-Guided Stereo Video Generation

2026-03-18 · Yang-Tian Sun, Zehuan Huang, Yifan Niu, Lin Ma 외 arxiv

We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates e…

Depth EstimationVideo Generation

Feature Learning for Interaction Activity Recognition in RGBD Videos

2015-08-10 · Ngu Nguyen

This paper proposes a human activity recognition method which is based on features learned from 3D video data without incorporating domain knowledge. The experiments on data collected by RGBD cameras produce results outp…

Activity RecognitionHuman Activity Recognition