paper-with-me

홈 › Papers

VIMI: Vehicle-Infrastructure Multi-view Intermediate Fusion for Camera-based 3D Object Detection

2023-03-20 · Zhe Wang, Siqi Fan, Xiaoliang Huo, Tongda Xu, Yan Wang, Jingjing Liu, Yilun Chen, Ya-Qin Zhang

In autonomous driving, Vehicle-Infrastructure Cooperative 3D Object Detection (VIC3D) makes use of multi-view cameras from both vehicles and traffic infrastructure, providing a global vantage point with rich semantic context of road conditions beyond a single vehicle viewpoint. Two major challenges prevail in VIC3D: 1) inherent calibration noise when fusing multi-view images, caused by time asynchrony across cameras; 2) information loss when projecting 2D features into 3D space. To address these issues, We propose a novel 3D object detection framework, Vehicles-Infrastructure Multi-view Intermediate fusion (VIMI). First, to fully exploit the holistic perspectives from both vehicles and infrastructure, we propose a Multi-scale Cross Attention (MCA) module that fuses infrastructure and vehicle features on selective multi-scales to correct the calibration noise introduced by camera asynchrony. Then, we design a Camera-aware Channel Masking (CCM) module that uses camera parameters as priors to augment the fused features. We further introduce a Feature Compression (FC) module with channel and spatial compression blocks to reduce the size of transmitted features for enhanced efficiency. Experiments show that VIMI achieves 15.61% overall AP_3D and 21.44% AP_BEV on the new VIC3D dataset, DAIR-V2X-C, significantly outperforming state-of-the-art early fusion and late fusion methods with comparable transmission cost.

📄 PDF Abstract BibTeX arXiv:2303.10975

Code (2)

bosszhe/vimi 공식 구현 pytorch
bosszhe/emiff pytorch

Tasks

3D Object DetectionAutonomous DrivingFeature Compressionobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Class-Adaptive Cooperative Perception for Multi-Class LiDAR-based 3D Object Detection in V2X Systems

2026-04-11 · Blessing Agyei Kyem, Joshua Kofi Asamoah, Armstrong Aboah arxiv

Cooperative perception allows connected vehicles and roadside infrastructure to share sensor observations, creating a fused scene representation beyond the capability of any single platform. However, most cooperative 3D …

3D Object Detection

VIMI: Grounding Video Generation through Multi-modal Instruction

2024-07-08 · Yuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen 외

Existing text-to-video diffusion models rely solely on text-only encoders for their pretraining. This limitation stems from the absence of large-scale multimodal prompt video datasets, resulting in a lack of visual groun…

Text-to-Video GenerationVideo GenerationVisual Grounding

I2V-GS: Infrastructure-to-Vehicle View Transformation with Gaussian Splatting for Autonomous Driving Data Generation

2025-07-31 · Jialei Chen, Wuhao Xu, Sipeng He, Baoru Huang 외 arxiv

Vast and high-quality data are essential for end-to-end autonomous driving systems. However, current driving data is mainly collected by vehicles, which is expensive and inefficient. A potential solution lies in synthesi…

Novel View SynthesisAutonomous Driving3D Reconstruction

EMIFF: Enhanced Multi-scale Image Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object Detection

2024-02-23 · Zhe Wang, Siqi Fan, Xiaoliang Huo, Tongda Xu 외

In autonomous driving, cooperative perception makes use of multi-view cameras from both vehicles and infrastructure, providing a global vantage point with rich semantic context of road conditions beyond a single vehicle …

3D Object DetectionAutonomous DrivingFeature Compressionobject-detection+1

ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access

2025-11-23 · Timing Yang, Sucheng Ren, Alan Yuille, Feng Wang arxiv

Text-to-video generation has surged in interest since Sora, yet open-source models still face a data bottleneck: there is no large, high-quality, easily obtainable video-text corpus. Existing public datasets typically re…

Video Question AnsweringText-to-Video Generation