paper-with-me

홈 › Papers

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding

2025-05-24 · Guofeng Mei, Bin Ren, Juan Liu, Luigi Riz, Xiaoshui Huang, Xu Zheng, Yongshun Gong, Ming-Hsuan Yang, Nicu Sebe, Fabio Poiesi

Vision-language models like CLIP can offer a promising foundation for 3D scene understanding when extended with 3D tokenizers. However, standard approaches, such as k-nearest neighbor or radius-based tokenization, struggle with cross-domain generalization due to sensitivity to dataset-specific spatial scales. We present a universal 3D tokenizer designed for scale-invariant representation learning with a frozen CLIP backbone. We show that combining superpoint-based grouping with coordinate scale normalization consistently outperforms conventional methods through extensive experimental analysis. Specifically, we introduce S4Token, a tokenization pipeline that produces semantically-informed tokens regardless of scene scale. Our tokenizer is trained without annotations using masked point modeling and clustering-based objectives, along with cross-modal distillation to align 3D tokens with 2D multi-view image features. For dense prediction tasks, we propose a superpoint-level feature propagation module to recover point-level detail from sparse tokens.

📄 PDF Abstract BibTeX arXiv:2505.18819

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationRepresentation LearningScene Understanding

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking

2019-10-28 · Alaaeldin El-Nouby, Shuangfei Zhai, Graham W. Taylor, Joshua M. Susskind

Deep neural networks require collecting and annotating large amounts of data to train successfully. In order to alleviate the annotation bottleneck, we propose a novel self-supervised representation learning approach for…

Action RecognitionFuture predictionRepresentation LearningSelf-Supervised Action Recognition

Self-supervised Discovery of Human Actons from Long Kinematic Videos

2021-09-29 · Kenneth Li, Xiao Sun, Zhirong Wu, Fangyun Wei 외

For human action understanding, a popular research direction is to analyze short video clips with unambiguous semantic content, such as jumping and drinking. However, methods for understanding short semantic actions cann…

Action UnderstandingSentence

Transfer CLIP for Generalizable Image Denoising

2024-03-22 · CVPR 2024 1 · Jun Cheng, Dong Liang, Shan Tan

Image denoising is a fundamental task in computer vision. While prevailing deep learning-based supervised and self-supervised methods have excelled in eliminating in-distribution noise, their susceptibility to out-of-dis…

DecoderDenoisingImage Denoising

Weakly-supervised HOI Detection via Prior-guided Bi-level Representation Learning

2023-03-02 · Bo Wan, Yongfei Liu, Desen Zhou, Tinne Tuytelaars 외

Human object interaction (HOI) detection plays a crucial role in human-centric scene understanding and serves as a fundamental building-block for many vision tasks. One generalizable and scalable strategy for HOI detecti…

Human-Object Interaction DetectionKnowledge DistillationObjectRepresentation Learning+1

Towards Tokenized Human Dynamics Representation

2021-11-22 · Kenneth Li, Xiao Sun, Zhirong Wu, Fangyun Wei 외

For human action understanding, a popular research direction is to analyze short video clips with unambiguous semantic content, such as jumping and drinking. However, methods for understanding short semantic actions cann…

Action SegmentationAction UnderstandingGenre classificationHuman Dynamics+1