paper-with-me

홈 › Papers

Training-free Spatially Grounded Geometric Shape Encoding (Technical Report)

2026-04-08 · Yuhang He arxiv

Positional encoding has become the de facto standard for grounding deep neural networks on discrete point-wise positions, and it has achieved remarkable success in tasks where the input can be represented as a one-dimensional sequence. However, extending this concept to 2D spatial geometric shapes demands carefully designed encoding strategies that account not only for shape geometry and pose, but also for compatibility with neural network learning. In this work, we address these challenges by introducing a training-free, general-purpose encoding strategy, dubbed XShapeEnc, that encodes an arbitrary spatially grounded 2D geometric shape into a compact representation exhibiting five favorable properties, including invertibility, adaptivity, and frequency richness. Specifically, a 2D spatially grounded geometric shape is decomposed into its normalized geometry within the unit disk and its pose vector, where the pose is further transformed into a harmonic pose field that also lies within the unit disk. A set of orthogonal Zernike bases is constructed to encode shape geometry and pose either independently or jointly, followed by a frequency-propagation operation to introduce high-frequency content into the encoding. We demonstrate the theoretical validity, efficiency, discriminability, and applicability of XShapeEnc via extensive analysis and experiments across a wide range of shape-aware tasks and our self-curated XShapeCorpus. We envision XShapeEnc as a foundational tool for research that goes beyond one-dimensional sequential data toward frontier 2D spatial intelligence.

📄 PDF Abstract BibTeX arXiv:2604.07522

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PARSE: Part-Aware Relational Spatial Modeling

2026-03-08 · Yinuo Bai, Peijun Xu, Kuixiang Shao, Yuyang Jiao 외 arxiv

Inter-object relations underpin spatial intelligence, yet existing representations -- linguistic prepositions or object-level scene graphs -- are too coarse to specify which regions actually support, contain, or contact …

Spatial Reasoning3D Generation

3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation

2025-06-11 · Seonho Lee, Jiho Choi, Inha Kang, Jiwook Kim 외

Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain fundamentally limited in their understanding of 3D spatial structures. We propose Geometric Distilla…

Spatial Reasoning

RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

2026-06-16 · Jinhao You, Shuo Lyu, Zhuohang Lyu, Tanxuan Li 외 arxiv

Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators re…

Simultaneous Monitoring of Shape and Surface Color via 4D Point Clouds: A Registration-free Approach

2026-05-09 · Mariafrancesca Patalano, Giovanna Capizzi, Kamran Paynabar arxiv

Advanced manufacturing technologies allow for the production of intricate parts featuring high shape complexity and spatially-varying material composition. Data fusion of point clouds with chromatic attributes provides 4…

Point Clouds

Ground4D: Spatially-Grounded Feedforward 4D Reconstruction for Unstructured Off-Road Scenes

2026-05-06 · Shuo Wang, Jilin Mei, Fuyang Liu, Wenfei Guan 외 arxiv

Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-road scenes, its performance degrades due to high-frequency geometry,…

Autonomous Driving