paper-with-me

Papers

Point-SRA: Self-Representation Alignment for 3D Representation Learning

2026-01-05 · Lintong Wei, Jian Lu, Haozhe Cheng, Jihua Zhu, Kaibing Zhang arxiv

Masked autoencoders (MAE) have become a dominant paradigm in 3D representation learning, setting new performance benchmarks across various downstream tasks. Existing methods with fixed mask ratio neglect multi-level representational correlations and intrinsic geometric structures, while relying on point-wise reconstruction assumptions that conflict with the diversity of point cloud. To address these issues, we propose a 3D representation learning method, termed Point-SRA, which aligns representations through self-distillation and probabilistic modeling. Specifically, we assign different masking ratios to the MAE to capture complementary geometric and semantic information, while the MeanFlow Transformer (MFT) leverages cross-modal conditional embeddings to enable diverse probabilistic reconstruction. Our analysis further reveals that representations at different time steps in MFT also exhibit complementarity. Therefore, a Dual Self-Representation Alignment mechanism is proposed at both the MAE and MFT levels. Finally, we design a Flow-Conditioned Fine-Tuning Architecture to fully exploit the point cloud distribution learned via MeanFlow. Point-SRA outperforms Point-MAE by 5.37% on ScanObjectNN. On intracranial aneurysm segmentation, it reaches 96.07% mean IoU for arteries and 86.87% for aneurysms. For 3D object detection, Point-SRA achieves 47.3% AP@50, surpassing MaskPoint by 5.12%.

📄 PDF Abstract BibTeX arXiv:2601.01746

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning3D Object Detection

Similar Papers 제목 키워드 기반

Self-supervised Human Mesh Recovery with Cross-Representation Alignment

2022-09-10 · Xuan Gong, Meng Zheng, Benjamin Planche, Srikrishna Karanam 외

Fully supervised human mesh recovery methods are data-hungry and have poor generalizability due to the limited availability and diversity of 3D-annotated benchmark datasets. Recent progress in self-supervised human mesh …

DiversityHuman Mesh Recovery

Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos

2023-08-18 · ICCV 2023 1 · Xiaoxiao Sheng, Zhiqiang Shen, Gang Xiao, Longguang Wang 외

We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at the clip or frame level and cannot well ca…

Contrastive LearningRepresentation LearningSelf-Supervised Learning

Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning

2025-09-29 · Donghwa Kang, Junho Kim, Dongwoo Kang arxiv

Event cameras offer unique advantages for facial keypoint alignment under challenging conditions, such as low light and rapid motion, due to their high temporal resolution and robustness to varying illumination. However,…

Representation Learning

How do Cross-View and Cross-Modal Alignment Affect Representations in Contrastive Learning?

2022-11-23 · Thomas M. Hehn, Julian F. P. Kooij, Dariu M. Gavrila

Various state-of-the-art self-supervised visual representation learning approaches take advantage of data from multiple sensors by aligning the feature representations across views and/or modalities. In this work, we inv…

Contrastive Learningcross-modal alignmentDepth EstimationDepth Prediction+5

PRNet: Self-Supervised Learning for Partial-to-Partial Registration

2019-10-27 · NeurIPS 2019 12 · Yue Wang, Justin M. Solomon

We present a simple, flexible, and general framework titled Partial Registration Network (PRNet), for partial-to-partial point cloud registration. Inspired by recently-proposed learning-based methods for registration, we…

Point Cloud RegistrationSelf-Supervised Learning