paper-with-me

홈 › Papers

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding

2025-08-07 · Changho Choi, Youngwoo Shin, Gyojin Han, Dong-Jae Lee, Junmo Kim arxiv

Understanding dynamic outdoor environments requires capturing complex object interactions and their evolution over time. LiDAR-based 4D point clouds provide precise spatial geometry and rich temporal cues, making them ideal for representing real-world scenes. However, despite their potential, 4D LiDAR remains underexplored in the context of Multimodal Large Language Models (MLLMs) due to the absence of high-quality, modality-specific annotations and the lack of MLLM architectures capable of processing its high-dimensional composition. To address these challenges, we introduce B4DL, a new benchmark specifically designed for training and evaluating MLLMs on 4D LiDAR understanding. In addition, we propose a scalable data generation pipeline and an MLLM model that, for the first time, directly processes raw 4D LiDAR by bridging it with language understanding. Combined with our dataset and benchmark, our model offers a unified solution for spatio-temporal reasoning in dynamic outdoor environments. We provide rendered 4D LiDAR videos, generated dataset, and inference outputs on diverse scenarios at: https://github.com/ccho4702/B4DL

📄 PDF Abstract BibTeX arXiv:2508.05269

Code (0)

등록된 구현이 없습니다.

Tasks

Point Clouds

Similar Papers 제목 키워드 기반

4D Panoptic LiDAR Segmentation

2021-02-24 · CVPR 2021 1 · Mehmet Aygün, Aljoša Ošep, Mark Weber, Maxim Maximov 외

Temporal semantic scene understanding is critical for self-driving cars or robots operating in dynamic environments. In this paper, we propose 4D panoptic LiDAR segmentation to assign a semantic class and a temporally-co…

4D Panoptic SegmentationBenchmarkingMulti-Object TrackingObject Tracking+3

SuperFlow++: Enhanced Spatiotemporal Consistency for Cross-Modal Data Pretraining

2025-03-25 · Xiang Xu, Lingdong Kong, Hui Shuai, Wenwei Zhang 외

LiDAR representation learning has emerged as a promising approach to reducing reliance on costly and labor-intensive human annotations. While existing methods primarily focus on spatial alignment between LiDAR and camera…

Autonomous DrivingComputational EfficiencyContrastive LearningRepresentation Learning+1

Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models

2026-08-26 · Jihun Kim, Hyun-Kurl Jang, Hyemin Yang, Jinnyeong Yang 외 arxiv

Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense …

Scene UnderstandingVideo Segmentation

A Spatiotemporal Correspondence Approach to Unsupervised LiDAR Segmentation with Traffic Applications

2023-08-23 · Xiao Li, Pan He, Aotian Wu, Sanjay Ranka 외

We address the problem of unsupervised semantic segmentation of outdoor LiDAR point clouds in diverse traffic scenarios. The key idea is to leverage the spatiotemporal nature of a dynamic point cloud sequence and introdu…

ClusteringPseudo LabelRepresentation LearningSegmentation+2

Joint 3D Object Detection and Tracking Using Spatio-Temporal Representation of Camera Image and LiDAR Point Clouds

2021-12-14 · Junho Koh, Jaekyum Kim, Jinhyuk Yoo, Yecheol Kim 외

In this paper, we propose a new joint object detection and tracking (JoDT) framework for 3D object detection and tracking based on camera and LiDAR sensors. The proposed method, referred to as 3D DetecTrack, enables the …

3D Object DetectionGraph Neural NetworkObjectobject-detection+1