paper-with-me

홈 › Papers

TexLiDAR: Automated Text Understanding for Panoramic LiDAR Data

2025-02-05 · Naor Cohen, Roy Orfaig, Ben-Zion Bobrovsky

Efforts to connect LiDAR data with text, such as LidarCLIP, have primarily focused on embedding 3D point clouds into CLIP text-image space. However, these approaches rely on 3D point clouds, which present challenges in encoding efficiency and neural network processing. With the advent of advanced LiDAR sensors like Ouster OS1, which, in addition to 3D point clouds, produce fixed resolution depth, signal, and ambient panoramic 2D images, new opportunities emerge for LiDAR based tasks. In this work, we propose an alternative approach to connect LiDAR data with text by leveraging 2D imagery generated by the OS1 sensor instead of 3D point clouds. Using the Florence 2 large model in a zero-shot setting, we perform image captioning and object detection. Our experiments demonstrate that Florence 2 generates more informative captions and achieves superior performance in object detection tasks compared to existing methods like CLIP. By combining advanced LiDAR sensor data with a large pre-trained model, our approach provides a robust and accurate solution for challenging detection scenarios, including real-time applications requiring high accuracy and robustness.

📄 PDF Abstract BibTeX arXiv:2502.04385

Code (1)

AIROTAU/TexLiDAR 공식 구현 pytorch

Tasks

Image Captioningobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Florence Florence is a computer vision foundation model aiming to learn universal visual-language representations that be adapted to various computer vision tasks, visual question…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Veila: Panoramic LiDAR Generation from a Monocular RGB Image

2025-08-05 · Youquan Liu, Lingdong Kong, Weidong Yang, Ao Liang 외 arxiv

Realistic and controllable panoramic LiDAR data generation is critical for scalable 3D perception in autonomous driving and robotics. Existing methods either perform unconditional generation with poor controllability or …

LIDAR Semantic SegmentationAutonomous DrivingData Augmentation

GS-LiDAR: Generating Realistic LiDAR Point Clouds with Panoramic Gaussian Splatting

2025-01-22 · Junzhe Jiang, Chun Gu, Yurui Chen, Li Zhang

LiDAR novel view synthesis (NVS) has emerged as a novel task within LiDAR simulation, offering valuable simulated point cloud data from novel viewpoints to aid in autonomous driving systems. However, existing LiDAR NVS m…

Autonomous DrivingNeRFNovel View Synthesis

Joint Calibration of Panoramic Camera and Lidar Based on Supervised Learning

2017-09-09 · Mingwei Cao, Ming Yang, Chunxiang Wang, Yeqiang Qian 외

In view of contemporary panoramic camera-laser scanner system, the traditional calibration method is not suitable for panoramic cameras whose imaging model is extremely nonlinear. The method based on statistical optimiza…

Translation

Multi-LVI-SAM: A Robust LiDAR-Visual-Inertial Odometry for Multiple Fisheye Cameras

2025-09-06 · Xinyu Zhang, Kai Huang, Junqiao Zhao, Zihan Yuan 외 arxiv

We propose a multi-camera LiDAR-visual-inertial odometry framework, Multi-LVI-SAM, which fuses data from multiple fisheye cameras, LiDAR and inertial sensors for highly accurate and robust state estimation. To enable eff…

Pose Estimation

HDPV-SLAM: Hybrid Depth-augmented Panoramic Visual SLAM for Mobile Mapping System with Tilted LiDAR and Panoramic Visual Camera

2023-01-27 · Mostafa Ahmadi, Amin Alizadeh Naeini, Mohammad Moein Sheikholeslami, Zahra Arjmandi 외

This paper proposes a novel visual simultaneous localization and mapping (SLAM) system called Hybrid Depth-augmented Panoramic Visual SLAM (HDPV-SLAM), that employs a panoramic camera and a tilted multi-beam LiDAR scanne…

Depth EstimationSimultaneous Localization and Mapping