Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
Vision-language models like CLIP can offer a promising foundation for 3D scene understanding when extended with 3D tokenizers. However, standard approaches, such as k-nearest neighbor or radius-based tokenization, struggle with cross-domain generalization due to sensitivity to dataset-specific spatial scales. We present a universal 3D tokenizer designed for scale-invariant representation learning with a frozen CLIP backbone. We show that combining superpoint-based grouping with coordinate scale normalization consistently outperforms conventional methods through extensive experimental analysis. Specifically, we introduce S4Token, a tokenization pipeline that produces semantically-informed tokens regardless of scene scale. Our tokenizer is trained without annotations using masked point modeling and clustering-based objectives, along with cross-modal distillation to align 3D tokens with 2D multi-view image features. For dense prediction tasks, we propose a superpoint-level feature propagation module to recover point-level detail from sparse tokens.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain GeneralizationRepresentation LearningScene UnderstandingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking
Deep neural networks require collecting and annotating large amounts of data to train successfully. In order to alleviate the annotation bottleneck, we propose a novel self-supervised representation learning approach for…
Action RecognitionFuture predictionRepresentation LearningSelf-Supervised Action RecognitionSelf-supervised Discovery of Human Actons from Long Kinematic Videos
For human action understanding, a popular research direction is to analyze short video clips with unambiguous semantic content, such as jumping and drinking. However, methods for understanding short semantic actions cann…
Action UnderstandingSentenceTransfer CLIP for Generalizable Image Denoising
Image denoising is a fundamental task in computer vision. While prevailing deep learning-based supervised and self-supervised methods have excelled in eliminating in-distribution noise, their susceptibility to out-of-dis…
DecoderDenoisingImage DenoisingWeakly-supervised HOI Detection via Prior-guided Bi-level Representation Learning
Human object interaction (HOI) detection plays a crucial role in human-centric scene understanding and serves as a fundamental building-block for many vision tasks. One generalizable and scalable strategy for HOI detecti…
Human-Object Interaction DetectionKnowledge DistillationObjectRepresentation Learning+1Towards Tokenized Human Dynamics Representation
For human action understanding, a popular research direction is to analyze short video clips with unambiguous semantic content, such as jumping and drinking. However, methods for understanding short semantic actions cann…
Action SegmentationAction UnderstandingGenre classificationHuman Dynamics+1