paper-with-me

홈 › Papers

Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving

2025-12-15 · Hyunki Seong, Jeong-Kyun Lee, Heesoo Myeong, Yongho Shin, Hyun-Mook Cho, Duck Hoon Kim, Pranav Desai, Monu Surana arxiv

Learning interactive motion behaviors among multiple agents is a core challenge in autonomous driving. While imitation learning models generate realistic trajectories, they often inherit biases from datasets dominated by safe demonstrations, limiting robustness in safety-critical cases. Moreover, most studies rely on open-loop evaluation, overlooking compounding errors in closed-loop execution. We address these limitations with two complementary strategies. First, we propose Group Relative Behavior Optimization (GRBO), a reinforcement learning post-training method that fine-tunes pretrained behavior models via group relative advantage maximization with human regularization. Using only 10% of the training dataset, GRBO improves safety performance by over 40% while preserving behavioral realism. Second, we introduce Warm-K, a warm-started Top-K sampling strategy that balances consistency and diversity in motion selection. Our Warm-K method-based test-time scaling enhances behavioral consistency and reactivity at test time without retraining, mitigating covariate shift and reducing performance discrepancies. Demo videos are available in the supplementary material.

📄 PDF Abstract BibTeX arXiv:2512.13262

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningAutonomous Driving

Similar Papers 제목 키워드 기반

Test-Time Scaling Makes Overtraining Compute-Optimal

2026-04-01 · Nicholas Roberts, Sungjun Cho, Zhiqi Gao, Tzu-Heng Huang 외 arxiv

Modern LLMs scale at test-time, e.g. via repeated sampling, where inference cost grows with model size and the number of samples. This creates a trade-off that pretraining scaling laws, such as Chinchilla, do not address…

Noise Conditional Variational Score Distillation

2025-06-11 · Xinyu Peng, Ziyang Zheng, Yaoming Wang, Han Li 외

We propose Noise Conditional Variational Score Distillation (NCVSD), a novel method for distilling pretrained diffusion models into generative denoisers. We achieve this by revealing that the unconditional score function…

Conditional Image GenerationDenoisingImage Generation

Sailing AI by the Stars: A Survey of Learning from Rewards in Post-Training and Test-Time Scaling of Large Language Models

2025-05-05 · Xiaobao Wu

Recent developments in Large Language Models (LLMs) have shifted from pre-training scaling to post-training and test-time scaling. Across these developments, a key unified paradigm has arisen: Learning from Rewards, wher…

Active Learning

Can Test-Time Scaling Improve World Foundation Model?

2025-03-31 · Wenyan Cong, Hanqing Zhu, Peihao Wang, Bangya Liu 외

World foundation models, which simulate the physical world by predicting future states from current observations and inputs, have become central to many applications in physical intelligence, including autonomous driving…

Autonomous Driving

Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models

2025-08-13 · Luca Eyring, Shyamgopal Karthik, Alexey Dosovitskiy, Nataniel Ruiz 외 arxiv

The new paradigm of test-time scaling has yielded remarkable breakthroughs in Large Language Models (LLMs) (e.g. reasoning models) and in generative vision models, allowing models to allocate additional computation durin…