paper-with-me

홈 › Papers

Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation

2026-07-01 · Zhaowen Zhu, Li Zhang, Yujie Chen, Tian Zhang, Yingjie Wang, Mingxia Zhan arxiv

Self-Supervised Monocular Depth Estimation (MDE) has garnered attention in recent years due to its independence from ground truth. However, most existing models are limited to a single scale and exhibit considerable performance degradation in complex driving environments. Networks specifically designed to handle dynamic traffic participants tend to be overly complex, hindering their deployment on resource-constrained automotive edge devices. To address these limitations and move towards robust driving perception, we propose FlexDepth, a scale-driven and flexible family of self-supervised MDE models tailored for challenging road scenarios. FlexDepth employs a two-stage static-dynamic decoupled training strategy, enabling the independent assessment of confidence for both static backgrounds and dynamic road objects. Furthermore, it introduces a meticulously designed Scale-Driven Decoder (SDD) to dynamically select components based on scale size, facilitating efficient feature fusion and the output of high-precision depth maps. Extensive experiments on standard driving benchmarks demonstrate that without any auxiliary information, our model achieves state-of-the-art performance across arbitrary scales with minimal computational overhead. Our smallest model, Flex-Nano, requires only 0.7 GFLOPs and achieves 37.6 FPS on mobile platforms, ensuring reliable real-time perception while maintaining excellent zero-shot generalization. Our source code is available: https://github.com/startnew/flexdepth

📄 PDF Abstract BibTeX arXiv:2607.00736

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationZero-shot Generalization

Similar Papers 제목 키워드 기반

DrivIng: A Large-Scale Multimodal Driving Dataset with Full Digital Twin Integration

2026-01-21 · Dominik Rößle, Xujun Xie, Adithya Mohan, Venkatesh Thirugnana Sambandham 외 arxiv

Perception is a cornerstone of autonomous driving, enabling vehicles to understand their surroundings and make safe, reliable decisions. Developing robust perception algorithms requires large-scale, high-quality datasets…

Autonomous Driving

X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability

2025-06-16 · Yu Yang, Alan Liang, Jianbiao Mei, Yukai Ma 외

Diffusion models are advancing autonomous driving by enabling realistic data synthesis, predictive end-to-end planning, and closed-loop simulation, with a primary focus on temporally consistent generation. However, the g…

3DGSAutonomous DrivingScene Generation

Prompt-Driven Domain Adaptation for End-to-End Autonomous Driving via In-Context RL

2025-11-16 · Aleesha Khurram, Amir Moeini, Shangtong Zhang, Rohan Chandra arxiv

Despite significant progress and advances in autonomous driving, many end-to-end systems still struggle with domain adaptation (DA), such as transferring a policy trained under clear weather to adverse weather conditions…

Reinforcement LearningAutonomous DrivingDomain Adaptation

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

2025-02-03 · Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng 외

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generatio…

Human AnimationHuman-Object Interaction DetectionMotion GenerationVideo Generation

Prioritizing Perception-Guided Self-Supervision: A New Paradigm for Causal Modeling in End-to-End Autonomous Driving

2025-11-11 · Yi Huang, Zhan Qu, Lihui Jiang, Bingbing Liu 외 arxiv

End-to-end autonomous driving systems, predominantly trained through imitation learning, have demonstrated considerable effectiveness in leveraging large-scale expert driving data. Despite their success in open-loop eval…

Autonomous Driving