paper-with-me

홈 › Papers

Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking

2025-10-23 · Zixuan Wu, Hengyuan Zhang, Ting-Hsuan Chen, Yuliang Guo, David Paz, Xinyu Huang, Liu Ren arxiv

Parking is a critical pillar of driving safety. While recent end-to-end (E2E) approaches have achieved promising in-domain results, robustness under domain shifts (e.g., weather and lighting changes) remains a key challenge. Rather than relying on additional data, in this paper, we propose Dino-Diffusion Parking (DDP), a domain-agnostic autonomous parking pipeline that integrates visual foundation models with diffusion-based planning to enable generalized perception and robust motion planning under distribution shifts. We train our pipeline in CARLA at regular setting and transfer it to more adversarial settings in a zero-shot fashion. Our model consistently achieves a parking success rate above 90% across all tested out-of-distribution (OOD) scenarios, with ablation studies confirming that both the network architecture and algorithmic design significantly enhance cross-domain performance over existing baselines. Furthermore, testing in a 3D Gaussian splatting (3DGS) environment reconstructed from a real-world parking lot demonstrates promising sim-to-real transfer.

📄 PDF Abstract BibTeX arXiv:2510.20335

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Planning

Similar Papers 제목 키워드 기반

ToMA: Token Merge with Attention for Diffusion Models

2025-09-13 · Wenbo Lu, Shaoyi Zheng, Yuxuan Xia, Shengjie Wang arxiv

Diffusion models excel in high-fidelity image generation but face scalability limits due to transformers' quadratic attention complexity. Plug-and-play token reduction methods like ToMeSD and ToFu reduce FLOPs by merging…

Image Generation

I Have an Attention Bridge to Sell You: Generalization Capabilities of Modular Translation Architectures

2024-04-27 · Timothee Mickus, Raúl Vázquez, Joseph Attieh

Modularity is a paradigm of machine translation with the potential of bringing forth models that are large at training time and small during inference. Within this field of study, modular approaches, and in particular at…

Machine TranslationTranslation

Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving

2026-02-05 · Xuyang Chen, Conglang Zhang, Chuanheng Fu, Zihao Yang 외 arxiv

Driven by the emergence of Controllable Video Diffusion, existing Sim2Real methods for autonomous driving video generation typically rely on explicit intermediate representations to bridge the domain gap. However, these …

Dimensionality ReductionAutonomous DrivingVideo Generation

P2P-Bridge: Diffusion Bridges for 3D Point Cloud Denoising

2024-08-29 · Mathias Vogel, Keisuke Tateno, Marc Pollefeys, Federico Tombari 외

In this work, we tackle the task of point cloud denoising through a novel framework that adapts Diffusion Schr\"odinger bridges to points clouds. Unlike previous approaches that predict point-wise displacements from poin…

Denoising

DINO-MX: A Modular & Flexible Framework for Self-Supervised Learning

2025-11-03 · Mahmut Selman Gokmen, Cody Bumgardner arxiv

Vision Foundation Models (VFMs) have advanced representation learning through self-supervised methods. However, existing training pipelines are often inflexible, domain-specific, or computationally expensive, which limit…

Self-Supervised LearningRepresentation LearningKnowledge DistillationData Augmentation