paper-with-me

홈 › Papers

DINO-SD: Champion Solution for ICRA 2024 RoboDepth Challenge

2024-05-27 · Yifan Mao, Ming Li, Jian Liu, Jiayang Liu, Zihan Qin, Chunxi Chu, Jialei Xu, Wenbo Zhao, Junjun Jiang, Xianming Liu

Surround-view depth estimation is a crucial task aims to acquire the depth maps of the surrounding views. It has many applications in real world scenarios such as autonomous driving, AR/VR and 3D reconstruction, etc. However, given that most of the data in the autonomous driving dataset is collected in daytime scenarios, this leads to poor depth model performance in the face of out-of-distribution(OoD) data. While some works try to improve the robustness of depth model under OoD data, these methods either require additional training data or lake generalizability. In this report, we introduce the DINO-SD, a novel surround-view depth estimation model. Our DINO-SD does not need additional data and has strong robustness. Our DINO-SD get the best performance in the track4 of ICRA 2024 RoboDepth Challenge.

📄 PDF Abstract BibTeX arXiv:2405.17102

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionAutonomous DrivingDepth Estimation

Similar Papers 제목 키워드 기반

Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Leveraging DINOv3 for Robust Outdoor Scene Understanding in Field Robotics

2026-06-17 · Jaeil Park, Hyobin Choi, Sangjin Lee, Hyungtae Lim 외 arxiv

The GOOSE 2D Fine-Grained Semantic Segmentation Challenge at the ICRA 2026 Workshop on Field Robotics evaluates dense semantic segmentation of off-road imagery over a fine-grained taxonomy of 64 classes and 11 evaluated …

Semantic SegmentationScene Understanding

Technical Report for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Pretraining-Diverse Ensemble of Foundation Vision Encoders for Robust Outdoor Scene Understanding

2026-06-22 · Boyan Wang, Yongxi Huang, Wenjing Li, Tianrui Hui 외 arxiv

This report presents our solution for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge, which requires parsing unstructured outdoor scenes from four camera platforms into 56 fine-grained categories. Ou…

Semantic SegmentationScene Understanding

The RoboDepth Challenge: Methods and Advancements Towards Robust Depth Estimation

2023-07-27 · Lingdong Kong, Yaru Niu, Shaoyuan Xie, Hanjiang Hu 외

Accurate depth estimation under out-of-distribution (OoD) scenarios, such as adverse weather conditions, sensor failure, and noise contamination, is desirable for safety-critical applications. Existing depth estimation s…

Depth EstimationImage RestorationSuper-Resolution

Taming VR Teleoperation and Learning from Demonstration for Multi-Task Bimanual Table Service Manipulation

2025-08-20 · Weize Li, Zhengxiao Han, Lixin Xu, Xiangyu Chen 외 arxiv

This technical report presents the champion solution of the Table Service Track in the ICRA 2025 What Bimanuals Can Do (WBCD) competition. We tackled a series of demanding tasks under strict requirements for speed, preci…

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

2025-08-04 · Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang 외 arxiv

Although Vision Language Models (VLMs) have shown strong generalization in medical imaging, pathology presents unique challenges due to ultra-high resolution, complex tissue structures, and nuanced clinical semantics. Th…

Visual Question AnsweringReinforcement Learning