paper-with-me

홈 › Papers

Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection

2026-03-23 · Kaiqiang Li, Gang Li, Mingle Zhou, Min Li, Delong Han, Jin Wan arxiv

Zero-shot (ZS) 3D anomaly detection is crucial for reliable industrial inspection, as it enables detecting and localizing defects without requiring any target-category training data. Existing approaches render 3D point clouds into 2D images and leverage pre-trained Vision-Language Models (VLMs) for anomaly detection. However, such strategies inevitably discard geometric details and exhibit limited sensitivity to local anomalies. In this paper, we revisit intrinsic 3D representations and explore the potential of pre-trained Point-Language Models (PLMs) for ZS 3D anomaly detection. We propose BTP (Back To Point), a novel framework that effectively aligns 3D point cloud and textual embeddings. Specifically, BTP aligns multi-granularity patch features with textual representations for localized anomaly detection, while incorporating geometric descriptors to enhance sensitivity to structural anomalies. Furthermore, we introduce a joint representation learning strategy that leverages auxiliary point cloud data to improve robustness and enrich anomaly semantics. Extensive experiments on Real3D-AD and Anomaly-ShapeNet demonstrate that BTP achieves superior performance in ZS 3D anomaly detection. Code will be available at \href{https://github.com/wistful-8029/BTP-3DAD}{https://github.com/wistful-8029/BTP-3DAD}.

📄 PDF Abstract BibTeX arXiv:2603.21511

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning3D Anomaly DetectionPoint Clouds

Similar Papers 제목 키워드 기반

SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation

2025-03-13 · Xiangyu Shi, Zerui Li, Wenqi Lyu, Jiatong Xia 외

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach…

Language ModelingLanguage ModellingLarge Language ModelVision and Language Navigation

A Mean Field Ansatz for Zero-Shot Weight Transfer

2024-08-16 · Xingyuan Chen, Wenwei Kuang, Lei Deng, Wei Han 외

The pre-training cost of large language models (LLMs) is prohibitive. One cutting-edge approach to reduce the cost is zero-shot weight transfer, also known as model growth for some cases, which magically transfers the we…

Safe Online Convex Optimization with Multi-Point Feedback

2024-07-16 · Spencer Hutchinson, Mahnoosh Alizadeh

Motivated by the stringent safety requirements that are often present in real-world applications, we study a safe online convex optimization setting where the player needs to simultaneously achieve sublinear regret and z…

Zero-shot point cloud segmentation by transferring geometric primitives

2022-10-18 · Runnan Chen, Xinge Zhu, Nenglun Chen, Wei Li 외

We investigate transductive zero-shot point cloud semantic segmentation, where the network is trained on seen objects and able to segment unseen objects. The 3D geometric elements are essential cues to imply a novel 3D o…

Point Cloud SegmentationSemantic Segmentation

More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery

2025-12-08 · Wenzhen Dong, Jieming Yu, Yiming Huang, Hongqiu Wang 외 arxiv

The recent SAM 3 and SAM 3D have introduced significant advancements over the predecessor, SAM 2, particularly with the integration of language-based segmentation and enhanced 3D perception capabilities. SAM 3 supports z…

Monocular Depth EstimationVideo Segmentation