paper-with-me

Papers

Multi-modal Streaming 3D Object Detection

2022-09-12 · Mazen Abdelfattah, Kaiwen Yuan, Z. Jane Wang, Rabab Ward

Modern autonomous vehicles rely heavily on mechanical LiDARs for perception. Current perception methods generally require 360{\deg} point clouds, collected sequentially as the LiDAR scans the azimuth and acquires consecutive wedge-shaped slices. The acquisition latency of a full scan (~ 100ms) may lead to outdated perception which is detrimental to safe operation. Recent streaming perception works proposed directly processing LiDAR slices and compensating for the narrow field of view (FOV) of a slice by reusing features from preceding slices. These works, however, are all based on a single modality and require past information which may be outdated. Meanwhile, images from high-frequency cameras can support streaming models as they provide a larger FoV compared to a LiDAR slice. However, this difference in FoV complicates sensor fusion. To address this research gap, we propose an innovative camera-LiDAR streaming 3D object detection framework that uses camera images instead of past LiDAR slices to provide an up-to-date, dense, and wide context for streaming perception. The proposed method outperforms prior streaming models on the challenging NuScenes benchmark. It also outperforms powerful full-scan detectors while being much faster. Our method is shown to be robust to missing camera images, narrow LiDAR slices, and small camera-LiDAR miscalibration.

📄 PDF Abstract BibTeX arXiv:2209.04966

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous VehiclesObjectobject-detectionObject DetectionSensor Fusion

Similar Papers 제목 키워드 기반

SODFormer: Streaming Object Detection with Transformer Using Events and Frames

2023-08-08 · Dianze Li, Jianing Li, Yonghong Tian

DAVIS camera, streaming two complementary sensing modalities of asynchronous events and frames, has gradually been used to address major object detection challenges (e.g., fast motion blur and low-light). However, how to…

object-detectionObject Detection

Streaming Object Detection for 3-D Point Clouds

2020-05-04 · ECCV 2020 8 · Wei Han, Zhengdong Zhang, Benjamin Caine, Brandon Yang 외

Autonomous vehicles operate in a dynamic environment, where the speed with which a vehicle can perceive and react impacts the safety and efficacy of the system. LiDAR provides a prominent sensory modality that informs ma…

Action RecognitionAutonomous VehiclesMotion EstimationObject+2

StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA

2025-10-29 · Yuhang Hu, Zhenyu Yang, Shihan Wang, Shengsheng Qian 외 arxiv

The rapid growth of streaming video applications demands multimodal models with enhanced capabilities for temporal dynamics understanding and complex reasoning. However, current Video Question Answering (VideoQA) dataset…

Video Question Answering

DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

2025-11-24 · Jiawei Hou, Shenghao Zhang, Can Wang, Zheng Gu 외 arxiv

Reliable 4D object detection, which refers to 3D object detection in streaming video, is crucial for perceiving and understanding the real world. Existing open-set 4D object detection methods typically make predictions o…

Multi-Task Learning3D Object Detection

A Multimodal Transformer for Live Streaming Highlight Prediction

2024-06-15 · Jiaxin Deng, Shiyao Wang, Dong Shen, Liqin Zhao 외

Recently, live streaming platforms have gained immense popularity. Traditional video highlight detection mainly focuses on visual features and utilizes both past and future content for prediction. However, live streaming…

Highlight DetectionPrediction