paper-with-me

홈 › Papers

MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning

2026-05-11 · Guangli Chen, Dianzhao Li, Wenjian Zhong, Bangquan Xie, Ostap Okhrin arxiv

Robust urban autonomous driving requires reliable 3D scene understanding and stable decision-making under dense interactions. However, existing end-to-end models lack interpretability, while modular pipelines suffer from error propagation across brittle interfaces. This paper proposes MTA-RL, the first framework that bridges perception and control through Multi-modal Transformer-based 3D Affordances and Reinforcement Learning (RL). Unlike previous fusion models that directly regress actions, RGB images and LiDAR point clouds are fused using a transformer architecture to predict explicit, geometry-aware affordance representations. These structured representations serve as a compact observation space, enabling the RL policy to operate purely on predicted driving semantics, which significantly improves sample efficiency and stability. Extensive evaluations in CARLA Town01-03 across varying densities (20-60 background vehicles) show that MTA-RL consistently outperforms state-of-the-art baselines. Trained solely on Town03, our method demonstrates superior zero-shot generalization in unseen towns, achieving up to a 9.0% increase in Route Completion, an 11.0% increase in Total Distance, and an 83.7% improvement in Distance Per Violation. Furthermore, ablation studies confirm that our multi-modal fusion and reward shaping are critical, significantly outperforming image-only and unshaped variants, demonstrating the effectiveness of MTA-RL for robust urban autonomous driving.

📄 PDF Abstract BibTeX arXiv:2605.10177

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationReinforcement LearningScene UnderstandingAutonomous Driving

Similar Papers 제목 키워드 기반

End-to-End Model-Free Reinforcement Learning for Urban Driving using Implicit Affordances

2019-11-25 · CVPR 2020 6 · Marin Toromanoff, Emilie Wirbel, Fabien Moutarde

Reinforcement Learning (RL) aims at learning an optimal behavior policy from its own experiments and not rule-based control methods. However, there is no RL algorithm yet capable of handling a task as difficult as urban …

Autonomous Drivingreinforcement-learningReinforcement LearningReinforcement Learning (RL)

MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer

2025-08-20 · Guile Wu, David Huang, Dongfeng Bai, Bingbing Liu arxiv

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack th…

Scene UnderstandingAutonomous DrivingVideo Generation

Learning End-to-end Autonomous Driving using Guided Auxiliary Supervision

2018-08-30 · Ashish Mehta, Adithya Subramanian, Anbumani Subramanian

Learning to drive faithfully in highly stochastic urban settings remains an open problem. To that end, we propose a Multi-task Learning from Demonstration (MT-LfD) framework which uses supervised auxiliary task predictio…

Autonomous DrivingMulti-Task Learning

Multi-Modal Fusion Transformer for End-to-End Autonomous Driving

2021-04-19 · CVPR 2021 1 · Aditya Prakash, Kashyap Chitta, Andreas Geiger

How should representations from complementary sensors be integrated for autonomous driving? Geometry-based sensor fusion has shown great promise for perception tasks such as object detection and motion forecasting. Howev…

Autonomous DrivingImitation LearningMotion Forecasting+4

End-To-End multi-modal sensors fusion system for urban automated driving

2018-10-10 · Ibrahim Sobh, Loay Amin, Sherif Abdelkarim, Khaled Elmadawy 외

In this paper, we present a novel framework for urban automated driving based on multi-modal sensors; LiDAR and Camera. Environment perception through sensors fusion is key to successful deployment of automated driving s…

SegmentationSemantic Segmentation