paper-with-me

Papers

ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation

2026-04-01 · Hao Zhang, Lue Fan, Weikang Bian, Zehuan Wu, Lewei Lu, Zhaoxiang Zhang, Hongsheng Li arxiv

We present ReinDriveGen, a framework that enables full controllability over dynamic driving scenes, allowing users to freely edit actor trajectories to simulate safety-critical corner cases such as front-vehicle collisions, drifting cars, vehicles spinning out of control, pedestrians jaywalking, and cyclists cutting across lanes. Our approach constructs a dynamic 3D point cloud scene from multi-frame LiDAR data, introduces a vehicle completion module to reconstruct full 360° geometry from partial observations, and renders the edited scene into 2D condition images that guide a video diffusion model to synthesize realistic driving videos. Since such edited scenarios inevitably fall outside the training distribution, we further propose an RL-based post-training strategy with a pairwise preference model and a pairwise reward mechanism, enabling robust quality improvement under out-of-distribution conditions without ground-truth supervision. Extensive experiments demonstrate that ReinDriveGen outperforms existing approaches on edited driving scenarios and achieves state-of-the-art results on novel ego viewpoint synthesis.

📄 PDF Abstract BibTeX arXiv:2604.01129

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Generation

Similar Papers 제목 키워드 기반

SCORP: Scene-Consistent Multi-agent Diffusion Planning with Stable Online Reinforcement Post-Training for Cooperative Driving

2026-04-13 · Haojie Bai, Aimin Li, Ruoyu Yao, Xiongwei Zhao 외 arxiv

Cooperative driving is a safety- and efficiency-critical task that requires the coordination of diverse, interaction-realistic multi-agent trajectories. Although existing diffusion-based methods can capture multimodal be…

Reinforcement Learning

Learning from Mistakes: Post-Training for Driving VLA with Takeover Data

2026-03-16 · Yinfeng Gao, Deqing Liu, Qichao Zhang, Yupeng Zheng 외 arxiv

Current Vision-Language-Action (VLA) paradigms in end-to-end autonomous driving rely on offline training from static datasets, leaving them vulnerable to distribution shift. Recent post-training methods use takeover data…

Autonomous Driving

World Engine: Towards the Era of Post-Training for Autonomous Driving

2026-06-18 · Tianyu Li, Li Chen, Caojun Wang, Haochen Liu 외 arxiv

Autonomous vehicles must operate safely in the real world, where errors can have severe consequences. Although modern end-to-end driving policies excel in routine scenarios, their reliability is limited by the scarcity o…

Autonomous VehiclesAutonomous Driving

CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies

2026-05-06 · Keyu Chen, Nanfei Ye, Yida Wang, Wenchao Sun 외 arxiv

Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit funda…

Autonomous Driving

Alignment Dynamics in LLM Fine-Tuning

2026-05-18 · Yuhan Huang, Huanran Chen, Yinpeng Dong arxiv

Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alignment is often fragile under subsequent fine-tuning. Existing explanat…

Reinforcement Learning