paper-with-me

Papers

Cross-Dimensional Refined Learning for Real-Time 3D Visual Perception from Monocular Video

2023-03-16 · Ziyang Hong, C. Patrick Yue

We present a novel real-time capable learning method that jointly perceives a 3D scene's geometry structure and semantic labels. Recent approaches to real-time 3D scene reconstruction mostly adopt a volumetric scheme, where a Truncated Signed Distance Function (TSDF) is directly regressed. However, these volumetric approaches tend to focus on the global coherence of their reconstructions, which leads to a lack of local geometric detail. To overcome this issue, we propose to leverage the latent geometric prior knowledge in 2D image features by explicit depth prediction and anchored feature generation, to refine the occupancy learning in TSDF volume. Besides, we find that this cross-dimensional feature refinement methodology can also be adopted for the semantic segmentation task by utilizing semantic priors. Hence, we proposed an end-to-end cross-dimensional refinement neural network (CDRNet) to extract both 3D mesh and 3D semantic labeling in real time. The experiment results show that this method achieves a state-of-the-art 3D perception efficiency on multiple datasets, which indicates the great potential of our method for industrial applications.

📄 PDF Abstract BibTeX arXiv:2303.09248

Code (0)

등록된 구현이 없습니다.

Tasks

3D Scene ReconstructionDepth EstimationDepth PredictionSemantic Segmentation

Similar Papers 제목 키워드 기반

ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models

2026-09-15 · Zhipeng Zhao, Zhaoqiang Wei, Peishun Liu, Youwei Zhao 외 arxiv

Gender bias in large vision-language models (LVLMs) undermines their fairness and reliability, compromising output trustworthiness. Current mitigation methods rely on training-phase adjustments or post-hoc calibration, b…

Visual GroundingImage CaptioningText Generation

Wide-Baseline Hair Capture Using Strand-Based Refinement

2013-06-01 · CVPR 2013 6 · Linjie Luo, Cha Zhang, Zhengyou Zhang, Szymon Rusinkiewicz

We propose a novel algorithm to reconstruct the 3D geometry of human hairs in wide-baseline setups using strand-based refinement. The hair strands are first extracted in each 2D view, and projected onto the 3D visual hul…

3D geometry

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance

2026-01-28 · Zhemeng Zhang, Jiahua Ma, Xincheng Yang, Xin Wen 외 arxiv

Fine-grained and contact-rich manipulation remain challenging for robots, largely due to the underutilization of tactile feedback. To address this, we introduce TouchGuide, a novel cross-policy visuo-tactile fusion parad…

Contrastive Learning

The Rashomon Effect for Visualizing High-Dimensional Data

2026-04-01 · Yiyang Sun, Haiyang Huang, Gaurav Rajesh Parikh, Cynthia Rudin arxiv

Dimension reduction (DR) is inherently non-unique: multiple embeddings can preserve the structure of high-dimensional data equally well while differing in layout or geometry. In this paper, we formally define the Rashomo…

Synthetic-to-Real Domain Bridging for Single-View 3D Reconstruction of Ships for Maritime Monitoring

2026-01-29 · Borja Carrillo-Perez, Felix Sattler, Angel Bueno Rodriguez, Maurice Stephan 외 arxiv

Three-dimensional (3D) reconstruction of ships is an important part of maritime monitoring, allowing improved visualization, inspection, and decision-making in real-world monitoring environments. However, most state-ofth…

Single-View 3D Reconstruction