paper-with-me

Papers

TriFusion-AE: Language-Guided Depth and LiDAR Fusion for Robust Point Cloud Processing

2025-09-23 · Susmit Neogi arxiv

LiDAR-based perception is central to autonomous driving and robotics, yet raw point clouds remain highly vulnerable to noise, occlusion, and adversarial corruptions. Autoencoders offer a natural framework for denoising and reconstruction, but their performance degrades under challenging real-world conditions. In this work, we propose TriFusion-AE, a multimodal cross-attention autoencoder that integrates textual priors, monocular depth maps from multi-view images, and LiDAR point clouds to improve robustness. By aligning semantic cues from text, geometric (depth) features from images, and spatial structure from LiDAR, TriFusion-AE learns representations that are resilient to stochastic noise and adversarial perturbations. Interestingly, while showing limited gains under mild perturbations, our model achieves significantly more robust reconstruction under strong adversarial attacks and heavy noise, where CNN-based autoencoders collapse. We evaluate on the nuScenes-mini dataset to reflect realistic low-data deployment scenarios. Our multimodal fusion framework is designed to be model-agnostic, enabling seamless integration with any CNN-based point cloud autoencoder for joint representation learning.

📄 PDF Abstract BibTeX arXiv:2509.18743

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningAutonomous DrivingPoint Clouds

Similar Papers 제목 키워드 기반

TriFusion-SR: Joint Tri-Modal Medical Image Fusion and SR

2026-03-10 · Fayaz Ali Dharejo, Sharif S. M. A., Aiman Khalil, Nachiket Chaudhary 외 arxiv

Multimodal medical image fusion facilitates comprehensive diagnosis by aggregating complementary structural and functional information, but its effectiveness is limited by resolution degradation and modality discrepancie…

Leveraging Sparse LiDAR for RAFT-Stereo: A Depth Pre-Fill Perspective

2025-07-26 · Jinsu Yoo, Sooyoung Jeon, Zanming Huang, Tai-Yu Pan 외 arxiv

We investigate LiDAR guidance within the RAFT-Stereo framework, aiming to improve stereo matching accuracy by injecting precise LiDAR depth into the initial disparity map. We find that the effectiveness of LiDAR guidance…

GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object Detection

2024-11-01 · CVPR 2024 1 · Xiaotian Li, Baojie Fan, Jiandong Tian, Huijie Fan

Recent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However, most of them overlook the complementary interaction and guidance …

3D Object Detectionobject-detectionObject Detection

DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception

2025-09-11 · Tim Broedermannn, Christos Sakaridis, Luigi Piccinelli, Wim Abbeloos 외 arxiv

Robust semantic perception for autonomous vehicles relies on effectively combining multiple sensors with complementary strengths and weaknesses. State-of-the-art sensor fusion approaches to semantic perception often trea…

Semantic SegmentationAutonomous Vehicles

SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction

2025-07-22 · Zaipeng Duan, Chenxu Dang, Xuzhong Hu, Pei An 외 arxiv

Multimodal 3D occupancy prediction has garnered significant attention for its potential in autonomous driving. However, most existing approaches are single-modality: camera-based methods lack depth information, while LiD…

Autonomous DrivingDepth Estimation