paper-with-me

홈 › Papers

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

2026-04-30 · Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, Dingyuan Zhang, Hengshuang Zhao, Xiang Bai arxiv

Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene generation, often overlooking comprehensive 3D scene understanding. Conversely, while Large Language Models (LLMs) demonstrate impressive reasoning capabilities, they lack the capacity to predict future geometric evolution, creating a significant disparity between semantic interpretation and physical simulation. To bridge this gap, we propose HERMES++, a unified driving world model that integrates 3D scene understanding and future geometry prediction within a single framework. Our approach addresses the distinct requirements of these tasks through synergistic designs. First, a BEV representation consolidates multi-view spatial information into a structure compatible with LLMs. Second, we introduce LLM-enhanced world queries to facilitate knowledge transfer from the understanding branch. Third, a Current-to-Future Link is designed to bridge the temporal gap, conditioning geometric evolution on semantic context. Finally, to enforce structural integrity, we employ a Joint Geometric Optimization strategy that integrates explicit geometric constraints with implicit latent regularization to align internal representations with geometry-aware priors. Extensive evaluations on multiple benchmarks validate the effectiveness of our method. HERMES++ achieves strong performance, outperforming specialist approaches in both future point cloud prediction and 3D scene understanding tasks. The model and code will be publicly released at https://github.com/H-EmbodVis/HERMESV2.

📄 PDF Abstract BibTeX arXiv:2604.28196

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingAutonomous DrivingScene Generation

Similar Papers 제목 키워드 기반

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

2025-01-24 · Xin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen 외

Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which invo…

Autonomous DrivingLanguage ModelingLanguage ModellingLarge Language Model+3

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

2026-02-01 · Weizhe Tang, Junwei You, Jiaxi Liu, Zhaoyi Wang 외 arxiv

End-to-end autonomous driving models increasingly benefit from large vision--language models for semantic understanding, yet ensuring safe and accurate operation under long-tail conditions remains challenging. These chal…

Trajectory PlanningAutonomous VehiclesAutonomous Driving

HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation

2025-02-17 · Ling Yang, Xinchen Zhang, Ye Tian, Chenming Shang 외

The remarkable success of the autoregressive paradigm has made significant advancement in Multimodal Large Language Models (MLLMs), with powerful models like Show-o, Transfusion and Emu3 achieving notable progress in uni…

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

2026-01-07 · Zhexiao Xiong, Xin Ye, Burhan Yaman, Sheng Cheng 외 arxiv

World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recent work has explored using vision-language models (VLMs) for planning, ye…

Scene UnderstandingTrajectory PlanningAutonomous DrivingImage Generation

DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding

2026-03-19 · Dong Zhuo, Wenzhao Zheng, Sicheng Zuo, Siming Yan 외 arxiv

With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visual modality. However, most existing toke…

Semantic SegmentationImage ReconstructionAutonomous Driving