paper-with-me

홈 › Papers

What Synthesis is Missing: Depth Adaptation Integrated with Weak Supervision for Indoor Scene Parsing

2019-03-23 · ICCV 2019 10 · Keng-Chi Liu, Yi-Ting Shen, Jan P. Klopp, Liang-Gee Chen

Scene Parsing is a crucial step to enable autonomous systems to understand and interact with their surroundings. Supervised deep learning methods have made great progress in solving scene parsing problems, however, come at the cost of laborious manual pixel-level annotation. To alleviate this effort synthetic data as well as weak supervision have both been investigated. Nonetheless, synthetically generated data still suffers from severe domain shift while weak labels are often imprecise. Moreover, most existing works for weakly supervised scene parsing are limited to salient foreground objects. The aim of this work is hence twofold: Exploit synthetic data where feasible and integrate weak supervision where necessary. More concretely, we address this goal by utilizing depth as transfer domain because its synthetic-to-real discrepancy is much lower than for color. At the same time, we perform weak localization from easily obtainable image level labels and integrate both using a novel contour-based scheme. Our approach is implemented as a teacher-student learning framework to solve the transfer learning problem by generating a pseudo ground truth. Using only depth-based adaptation, this approach already outperforms previous transfer learning approaches on the popular indoor scene parsing SUN RGB-D dataset. Our proposed two-stage integration more than halves the gap towards fully supervised methods when compared to previous state-of-the-art in transfer learning.

📄 PDF Abstract BibTeX arXiv:1903.09781

Code (0)

등록된 구현이 없습니다.

Tasks

Scene ParsingTransfer Learning

Similar Papers 제목 키워드 기반

GeomPrompt: Geometric Prompt Learning for RGB-D Semantic Segmentation Under Missing and Degraded Depth

2026-04-13 · Krishna Jaganathan, Patricio Vela arxiv

Multimodal perception systems for robotics and embodied AI often assume reliable RGB-D sensing, but in practice, depth is frequently missing, noisy, or corrupted. We thus present GeomPrompt, a lightweight cross-modal ada…

Semantic Segmentation

DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving

2025-09-09 · Sven Kirchner, Nils Purschke, Ross Greer, Alois C. Knoll arxiv

Ensuring reliable autonomous operation when visual input is degraded remains a key challenge in intelligent vehicles and robotics. We present DepthVision, a multimodal framework that enables Vision--Language Models (VLMs…

Scene UnderstandingAutonomous DrivingPoint Clouds

Domain-Adaptive 3D Medical Image Synthesis: An Efficient Unsupervised Approach

2022-07-02 · Qingqiao Hu, Hongwei Li, JianGuo Zhang

Medical image synthesis has attracted increasing attention because it could generate missing image data, improving diagnosis and benefits many downstream tasks. However, so far the developed synthesis model is not adapti…

Domain AdaptationImage Generation

ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative Adaptation

2026-01-15 · Kim Youwang, Lee Hyoseok, Subin Park, Gerard Pons-Moll 외 arxiv

We introduce ELITE, an Efficient Gaussian head avatar synthesis from a monocular video via Learned Initialization and TEst-time generative adaptation. Prior works rely either on a 3D data prior or a 2D generative prior t…

Deep Learning-based High-precision Depth Map Estimation from Missing Viewpoints for 360 Degree Digital Holography

2021-03-09 · Hakdong Kim, Heonyeong Lim, Minkyu Jee, Yurim Lee 외

In this paper, we propose a novel, convolutional neural network model to extract highly precise depth maps from missing viewpoints, especially well applicable to generate holographic 3D contents. The depth map is an esse…