paper-with-me

홈 › Papers

Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds

2025-12-09 · Shaofeng Zhang, Xuanqi Chen, Xiangdong Zhang, Sitong Wu, Junchi Yan arxiv

Most existing self-supervised learning (SSL) approaches for 3D point clouds are dominated by generative methods based on Masked Autoencoders (MAE). However, these generative methods have been proven to struggle to capture high-level discriminative features effectively, leading to poor performance on linear probing and other downstream tasks. In contrast, contrastive methods excel in discriminative feature representation and generalization ability on image data. Despite this, contrastive learning (CL) in 3D data remains scarce. Besides, simply applying CL methods designed for 2D data to 3D fails to effectively learn 3D local details. To address these challenges, we propose a novel Dual-Branch \textbf{C}enter-\textbf{S}urrounding \textbf{Con}trast (CSCon) framework. Specifically, we apply masking to the center and surrounding parts separately, constructing dual-branch inputs with center-biased and surrounding-biased representations to better capture rich geometric information. Meanwhile, we introduce a patch-level contrastive loss to further enhance both high-level information and local sensitivity. Under the FULL and ALL protocols, CSCon achieves performance comparable to generative methods; under the MLP-LINEAR, MLP-3, and ONLY-NEW protocols, our method attains state-of-the-art results, even surpassing cross-modal approaches. In particular, under the MLP-LINEAR protocol, our method outperforms the baseline (Point-MAE) by \textbf{7.9\%}, \textbf{6.7\%}, and \textbf{10.3\%} on the three variants of ScanObjectNN, respectively. The code will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2512.08673

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningContrastive LearningPoint Clouds

Similar Papers 제목 키워드 기반

PD-APE: A Parallel Decoding Framework with Adaptive Position Encoding for 3D Visual Grounding

2024-07-19 · Chenshu Hou, Liang Peng, Xiaopei Wu, Xiaofei He 외

3D visual grounding aims to identify objects in 3D point cloud scenes that match specific natural language descriptions. This requires the model to not only focus on the target object itself but also to consider the surr…

3D visual groundingAttributeDecoderLanguage Modelling+3

Dual-branch residual network for lung nodule segmentation

2019-05-21 · Haichao Cao, Hong Liu, Enmin Song, Chih-Cheng Hung 외

An accurate segmentation of lung nodules in computed tomography (CT) images is critical to lung cancer analysis and diagnosis. However, due to the variety of lung nodules and the similarity of visual characteristics betw…

Computed Tomography (CT)Lung Nodule SegmentationSegmentation

SCINet: Spatial and Contrast Interactive Super-Resolution Assisted Infrared UAV Target Detection

2024-10-01 · IEEE Transactions on Geoscience and Remote Sensing 2024 10 · Houzhang Fang, Lan Ding, Xiaolin Wang, Yi Chang 외

Unmanned aerial vehicle (UAV) detection based on thermal infrared imaging has been one of the most important sensing technologies in the anti-UAV system. However, the technical limitations and long-range detection of …

Super-Resolution

Rethinking Explanations: Formalizing Contrast in Description Logics

2026-05-02 · Yasir Mahmood, Arnab Sharma, Axel-Cyrille Ngonga Ngomo, Balram Tiwari arxiv

There has been a growing interest in explaining entailments over description logic (DL) knowledge bases. The existing explanation formalisms focus on justifications to explain true axioms, and abductive reasoning to expl…

VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models

2024-10-02 · Kailai Feng, Yabo Zhang, Haodong Yu, Zhilong Ji 외

Artistic typography is a technique to visualize the meaning of input character in an imaginable and readable manner. With powerful text-to-image diffusion models, existing methods directly design the overall geometry and…