OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs
Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Autonomous Vehicles (AV) remains an open challenge. Existing geometry-aware MLLMs typically rely on auxiliary 3D models at inference time, introducing pipeline complexity and the risk of cascading failures. In this paper, we present OmniSpace, a simple yet effective plug-and-play paradigm for geometry-aware spatial reasoning from purely 2D observations. Motivated by our finding that current MLLMs are bottlenecked by weak cross-view correspondence and depth estimation, OmniSpace introduces a Camera Pose Injector, a Multi-view Epipolar Attention module, and a 3D Geometric Distillation objective that jointly address these two limitations by transferring geometric knowledge into the model. Extensive experiments show that OmniSpace surpasses existing methods on planning benchmarks (nuScenes, Bench2Drive), risk detection (nuInstruct), language (Omnidrive), and generalization (DriveBench).
Code (0)
등록된 구현이 없습니다.
Tasks
Autonomous VehiclesSpatial ReasoningDepth EstimationSimilar Papers 제목 키워드 기반
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understanding of Multimodal Large Language Model…
Multimodal ReasoningSpatial ReasoningVisual ReasoningIncorporating Explanations into Human-Machine Interfaces for Trust and Situation Awareness in Autonomous Vehicles
Autonomous vehicles often make complex decisions via machine learning-based predictive models applied to collected sensor data. While this combination of methods provides a foundation for real-time actions, self-driving …
Autonomous VehiclesScene UnderstandingLearning Multi-Modal Self-Awareness Models for Autonomous Vehicles from Human Driving
This paper presents a novel approach for learning self-awareness models for autonomous vehicles. The proposed technique is based on the availability of synchronized multi-sensor dynamic data related to different maneuver…
Anomaly DetectionAutonomous VehiclesDecision MakingSelf-awareness in intelligent vehicles: Feature based dynamic Bayesian models for abnormality detection
The evolution of Intelligent Transportation Systems in recent times necessitates the development of self-awareness in agents. Before the intensive use of Machine Learning, the detection of abnormalities was manually prog…
Anomaly DetectionAutonomous VehiclesTime SeriesTime Series AnalysisGeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
Recent advancements in world models have revolutionized dynamic environment simulation, allowing systems to foresee future states and assess potential actions. In autonomous driving, these capabilities help vehicles anti…
3D geometryAutonomous DrivingAutonomous NavigationOcclusion Handling