paper-with-me

홈 › Papers

Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving

2026-03-06 · Yuhan Zhou, Mehri Sattari, Haihua Chen, Kewei Sha arxiv

Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-making. In practice, data quality (DQ) varies across sources and modalities due to environmental conditions and sensor limitations, yet AV research has largely prioritized algorithm design over DQ analysis. This work focuses on redundancy as a fundamental but underexplored DQ issue in AV datasets. Using the nuScenes and Argoverse 2 (AV2) datasets, we model and measure redundancy in multisource camera data and multimodal image-LiDAR data, and evaluate how removing redundant labels affects the YOLOv8 object detection task. Experimental results show that selectively removing redundant multisource image object labels from cameras with shared fields of view improves detection. In nuScenes, mAP${50}$ gains from $0.66$ to $0.70$, $0.64$ to $0.67$, and from $0.53$ to $0.55$, on three representative overlap regions, while detection on other overlapping camera pairs remains at the baseline even under stronger pruning. In AV2, $4.1$-$8.6\%$ of labels are removed, and mAP${50}$ stays near the $0.64$ baseline. Multimodal analysis also reveals substantial redundancy between image and LiDAR data. These findings demonstrate that redundancy is a measurable and actionable DQ factor with direct implications for AV performance. This work highlights the role of redundancy as a data quality factor in AV perception and motivates a data-centric perspective for evaluating and improving AV datasets. Code, data, and implementation details are publicly available at: https://github.com/yhZHOU515/RedundancyAD

📄 PDF Abstract BibTeX arXiv:2603.06544

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous VehiclesAutonomous DrivingObject Detection

Similar Papers 제목 키워드 기반

Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain Videos

2020-11-01 · EMNLP 2020 11 · Nayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang 외

Multimodal summarization for open-domain videos is an emerging task, aiming to generate a summary from multisource information (video, audio, transcript). Despite the success of recent multiencoder-decoder frameworks on …

Decoder

Multisource and Multitemporal Data Fusion in Remote Sensing

2018-12-19 · Pedram Ghamisi, Behnood Rasti, Naoto Yokoya, Qunming Wang 외

The sharp and recent increase in the availability of data captured by different sensors combined with their considerably heterogeneous natures poses a serious challenge for the effective and efficient processing of remot…

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

2026-06-23 · Bin Chen, Yuxiang Cai, Yadan Luo, Yi Zhang 외 arxiv

Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performance. Existing token pruning methods typically rely on single-layer si…

APT-MMF: An advanced persistent threat actor attribution method based on multimodal and multilevel feature fusion

2024-02-20 · Nan Xiao, Bo Lang, Ting Wang, Yikai Chen

Threat actor attribution is a crucial defense strategy for combating advanced persistent threats (APTs). Cyber threat intelligence (CTI), which involves analyzing multisource heterogeneous data from APTs, plays an import…

AttributeGraph Attention

OVERLORD: Ultimate Scaling of DataLoader for Multi-Source Large Foundation Model Training

2025-04-14 · Juntao Zhao, Qi Lu, Wei Jia, Borui Wan 외

Modern frameworks for training large foundation models (LFMs) employ dataloaders in a data-parallel manner, with each loader processing a disjoint subset of training data. Under multisource preprocessing, two fundamental…

CPU