paper-with-me

홈 › Papers

DeepIPCv3: Event-Aware Multi-Modal Sensor Fusion for Sudden Pedestrian Crossing Avoidance

2026-05-31 · Oskar Natan, Andi Dharmawan, Aufaclav Zatu Kusuma Frisky, Jazi Eko Istiyanto, Jun Miura arxiv

Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur during highly dynamic encounters, specifically sudden pedestrian crossings. To address this critical safety vulnerability, we propose DeepIPCv3, a novel multi-modal autonomous navigation framework that synergizes the dense 3D spatial geometry of LiDAR point clouds with the microsecond-level asynchronous event streams of a Dynamic Vision Sensor (DVS). We introduce a Transformer-inspired cross-modal attention mechanism to dynamically correlate these distinct modalities, allowing the network to instantaneously prioritize high-speed dynamic updates without sacrificing structural scene awareness. The fused latent representations are then mapped to safe local waypoints and executable control commands via a hybrid policy network that blends heuristic trajectory tracking with direct neural predictions. Due to the severe physical risks associated with live testing of these sudden crossing scenarios, the framework is rigorously evaluated offline using a custom multi-modal dataset collected across both well-illuminated noon and challenging evening conditions. Extensive comparative and ablation studies demonstrate that DeepIPCv3 achieves state-of-the-art predictive performance. By effectively eliminating exposure failures and motion blur, the proposed LiDAR and DVS fusion yields the lowest trajectory and control command errors, enabling highly reactive, mathematically bounded evasive maneuvers regardless of ambient illumination. To support future research, we will release the codes to our GitHub repo at https://github.com/oskarnatan/DeepIPCv3.

📄 PDF Abstract BibTeX arXiv:2606.01277

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingPoint Clouds

Similar Papers 제목 키워드 기반

DeepIPCv2: LiDAR-powered Robust Environmental Perception and Navigational Control for Autonomous Vehicle

2023-07-13 · Oskar Natan, Jun Miura

We present DeepIPCv2, an autonomous driving model that perceives the environment using a LiDAR sensor for more robust drivability, especially when driving under poor illumination conditions where everything is not clearl…

Autonomous DrivingScene Understanding

Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection

2025-05-05 · Sungheon Jeong, Jihong Park, Mohsen Imani

Most existing video anomaly detectors rely solely on RGB frames, which lack the temporal resolution needed to capture abrupt or transient motion cues, key indicators of anomalous events. To address this limitation, we pr…

Anomaly DetectionAnomaly Detection In Surveillance VideosVideo Anomaly DetectionVideo Understanding

Learning Multi-Modal Self-Awareness Models for Autonomous Vehicles from Human Driving

2018-06-07 · Mahdyar Ravanbakhsh, Mohamad Baydoun, Damian Campo, Pablo Marin 외

This paper presents a novel approach for learning self-awareness models for autonomous vehicles. The proposed technique is based on the availability of synchronized multi-sensor dynamic data related to different maneuver…

Anomaly DetectionAutonomous VehiclesDecision Making

Multi-view and Multi-modal Event Detection Utilizing Transformer-based Multi-sensor fusion

2022-02-18 · Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito, Noboru Harada

We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microphones and their weak labels. In this task…

Event DetectionSensor Fusion

Exploring Context, Attention and Audio Features for Audio Visual Scene-Aware Dialog

2019-12-20 · Shachi H. Kumar, Eda Okur, Saurav Sahay, Jonathan Huang 외

We are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs to learn audio-visual groundings of utterances and have conversations with users about the objects, activities an…

Audio ClassificationVisual Grounding