paper-with-me

홈 › Papers

A Multimodal Hybrid Late-Cascade Fusion Network for Enhanced 3D Object Detection

2025-04-25 · Carlo Sgaravatti, Roberto Basla, Riccardo Pieroni, Matteo Corno, Sergio M. Savaresi, Luca Magri, Giacomo Boracchi

We present a new way to detect 3D objects from multimodal inputs, leveraging both LiDAR and RGB cameras in a hybrid late-cascade scheme, that combines an RGB detection network and a 3D LiDAR detector. We exploit late fusion principles to reduce LiDAR False Positives, matching LiDAR detections with RGB ones by projecting the LiDAR bounding boxes on the image. We rely on cascade fusion principles to recover LiDAR False Negatives leveraging epipolar constraints and frustums generated by RGB detections of separate views. Our solution can be plugged on top of any underlying single-modal detectors, enabling a flexible training process that can take advantage of pre-trained LiDAR and RGB detectors, or train the two branches separately. We evaluate our results on the KITTI object detection benchmark, showing significant performance improvements, especially for the detection of Pedestrians and Cyclists.

📄 PDF Abstract BibTeX arXiv:2504.18419

Code (1)

CarloSgaravatti/HybridLateCascadeFusion 공식 구현 pytorch

Tasks

3D Object Detectionobject-detectionObject Detection

Similar Papers 제목 키워드 기반

HyPCA-Net: Advancing Multimodal Fusion in Medical Image Analysis

2026-02-18 · J. Dhar, M. K. Pandey, D. Chakladar, M. Haghighat 외 arxiv

Multimodal fusion frameworks, which integrate diverse medical imaging modalities (e.g., MRI, CT), have shown great potential in applications such as skin cancer detection, dementia diagnosis, and brain tumor prediction. …

Multimodal Priors-Augmented Text-Driven 3D Human-Object Interaction Generation

2026-02-11 · Yin Wang, Ziyao Zhang, Zhiying Leng, Haitian Liu 외 arxiv

We address the challenging task of text-driven 3D human-object interaction (HOI) motion generation. Existing methods primarily rely on a direct text-to-HOI mapping, which suffers from three key limitations due to the sig…

Leveraging Cascaded Binary Classification and Multimodal Fusion for Dementia Detection through Spontaneous Speech

2025-05-26 · Yin-Long Liu, Yuanchao Li, Rui Feng, Liu He 외

This paper presents our submission to the PROCESS Challenge 2025, focusing on spontaneous speech analysis for early dementia detection. For the three-class classification task (Healthy Control, Mild Cognitive Impairment,…

Binary ClassificationClassificationEnsemble LearningMulti-class Classification+1

Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions

2026-06-08 · Luxury, Jie Huang, Zihao Fan, Xiaoxiao Ma 외 arxiv

While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving efficient, scalable, real-time high-resolution video generation a fun…

Video Generation

Cross-Platform E-Commerce Product Categorization and Recategorization: A Multimodal Hierarchical Classification Approach

2025-08-27 · Lotte Gross, Rebecca Walter, Nicole Zoppi, Adrien Justus 외 arxiv

This study addresses critical industrial challenges in e-commerce product categorization, namely platform heterogeneity and the structural limitations of existing taxonomies, by developing and deploying a multimodal hier…

Product Categorization