paper-with-me

홈 › Papers

Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution

2021-03-24 · CVPR 2021 1 · Baoli Sun, Xinchen Ye, Baopu Li, Haojie Li, Zhihui Wang, Rui Xu

Existing color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training samples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected in actual testing environment. Therefore, we explore for the first time to learn the cross-modality knowledge at training stage, where both RGB and depth modalities are available, but test on the target dataset, where only single depth modality exists. Our key idea is to distill the knowledge of scene structural guidance from RGB modality to the single DSR task without changing its network architecture. Specifically, we construct an auxiliary depth estimation (DE) task that takes an RGB image as input to estimate a depth map, and train both DSR task and DE task collaboratively to boost the performance of DSR. Upon this, a cross-task interaction module is proposed to realize bilateral cross task knowledge transfer. First, we design a cross-task distillation scheme that encourages DSR and DE networks to learn from each other in a teacher-student role-exchanging fashion. Then, we advance a structure prediction (SP) task that provides extra structure regularization to help both DSR and DE networks learn more informative structure representations for depth recovery. Extensive experiments demonstrate that our scheme achieves superior performance in comparison with other DSR methods.

📄 PDF Abstract BibTeX arXiv:2103.12955

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationSuper-ResolutionTransfer Learning

Similar Papers 제목 키워드 기반

VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion

2025-03-08 · Meng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin 외

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity cau…

3D Semantic Scene CompletionAutonomous DrivingLanguage ModelingLanguage Modelling+1

S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance

2025-12-01 · Beining Xu, Siting Zhu, Zhao Jin, Junxian Li 외 arxiv

3D Visual Grounding (3DVG) focuses on locating objects in 3D scenes based on natural language descriptions, serving as a fundamental task for embodied AI and robotics. Recent advances in Multi-modal Large Language Models…

Spatial Reasoning3D ReconstructionVisual GroundingPoint Clouds

Task-specific Scene Structure Representations

2023-01-02 · Jisu Shin, Seunghyun Shin, Hae-Gon Jeon

Understanding the informative structures of scenes is essential for low-level vision tasks. Unfortunately, it is difficult to obtain a concrete visual definition of the informative structures because influences of visual…

Denoisinggraph partitioningImage Denoising

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

2025-05-05 · Lu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding 외

Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several challenges. Learning-based approaches depend on small-scale indoor datasets, …

Common Sense ReasoningScene Generation

Semantically Adversarial Scenario Generation with Explicit Knowledge Guidance

2021-06-08 · Wenhao Ding, Haohong Lin, Bo Li, Ding Zhao

Generating adversarial scenarios, which have the potential to fail autonomous driving systems, provides an effective way to improve robustness. Extending purely data-driven generative models, recent specialized models sa…

Autonomous DrivingAutonomous VehiclesPoint Cloud SegmentationScene Generation