paper-with-me

홈 › Papers

TOLiD: Bridging the Architecture Gap in Vision Foundation Model to LiDAR Pretraining via Token Lifting for Distillation

2026-07-12 · Sutharsan Mahendran, Darshana Priyasad, Kaushik Roy, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Peyman Moghadam arxiv

Cross-modal distillation from Vision Foundation Models (VFMs) to LiDAR backbones has recently emerged as a self-supervised pretraining strategy that reduces reliance on dense point-wise annotation for 3D scene understanding. However, existing distillation pipelines typically treat the VFM as a frozen feature source and train a heterogeneous 3D backbone to match fixed image embeddings, forcing the student to bridge both the modality gap and the cross-architecture gap between dense ViT token representations and sparse 3D encoders. We propose TOLiD, a self-supervised pretraining method for LiDAR representation learning that addresses this gap by coupling a LiDAR backbone with a student Vision Transformer (ViT) initialized from a frozen VFM teacher and applying supervision over compatible patch-token representations. TOLiD converts the set of point features within each image patch frustum into a token using Frustum Pooling followed by Frustum Attention, and performs token-level distillation with visibility masking. For LiDAR-only deployment, we lift token features back to per-point representations using masked bilinear sampling to avoid patches that have limited LiDAR points. We extensively evaluate TOLiD on five heterogeneous LiDAR datasets and four cross-sensor adaptation pairs, demonstrating improved transfer with frozen backbones and lightweight heads.

📄 PDF Abstract BibTeX arXiv:2607.10762

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningScene Understanding

Similar Papers 제목 키워드 기반

Label-Efficient LiDAR Semantic Segmentation with 2D-3D Vision Transformer Adapters

2025-03-05 · Julia Hindel, Rohit Mohan, Jelena Bratulic, Daniele Cattaneo 외

LiDAR semantic segmentation models are typically trained from random initialization as universal pre-training is hindered by the lack of large, diverse datasets. Moreover, most point cloud segmentation architectures inco…

LIDAR Semantic SegmentationPoint Cloud SegmentationSemantic Segmentation

Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception

2026-04-19 · Siyuan Meng, Chengbo Ai arxiv

World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego-vehicle perspective, leaving the infrastructure viewpoint unexplore…

Scene UnderstandingAutonomous Driving

Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift

2025-11-21 · Björn Michele, Alexandre Boulch, Gilles Puy, Tuan-Hung Vu 외 arxiv

Semantic segmentation networks trained under full supervision for one type of lidar fail to generalize to unseen lidars without intervention. To reduce the performance gap under domain shifts, a recent trend is to levera…

Unsupervised Domain Adaptation3D Semantic SegmentationKnowledge DistillationPoint Clouds

Better Call SAL: Towards Learning to Segment Anything in Lidar

2024-03-19 · Aljoša Ošep, Tim Meinhardt, Francesco Ferroni, Neehar Peri 외

We propose the SAL (Segment Anything in Lidar) method consisting of a text-promptable zero-shot model for segmenting and classifying any object in Lidar, and a pseudo-labeling engine that facilitates model training witho…

Panoptic SegmentationSegmentation

Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models

2026-08-26 · Jihun Kim, Hyun-Kurl Jang, Hyemin Yang, Jinnyeong Yang 외 arxiv

Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense …

Scene UnderstandingVideo Segmentation