paper-with-me

홈 › Papers

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization

2026-03-09 · Bryce Grant, Aryeh Rothenberg, Atri Banerjee, Peng Wang arxiv

Localizing objects and parts from natural language in 3D space is essential for robotics, AR, and embodied AI, yet existing methods face a trade-off between the accuracy and geometric consistency of per-scene optimization and the efficiency of feed-forward inference. We present TrianguLang, a feed-forward framework for 3D localization that requires no camera calibration at inference. Unlike prior methods that treat views independently, we introduce Geometry-Aware Semantic Attention (GASA), which utilizes predicted geometry to gate cross-view feature correspondence, suppressing semantically-plausible but geometrically-inconsistent matches without requiring ground-truth poses. Validated on five benchmarks including ScanNet++ and uCO3D, TrianguLang achieves state-of-the-art feed-forward text-guided segmentation and localization, reducing user effort from $O(N)$ clicks to a single text query. The model processes each frame at 1008x1008 resolution in $\sim$57ms ($\sim$18 FPS) without optimization, enabling practical deployment for interactive robotics and AR applications. Code and checkpoints are available at https://cwru-aism.github.io/triangulang/.

📄 PDF Abstract BibTeX arXiv:2603.08096

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CoRegOVCD: Consistency-Regularized Open-Vocabulary Change Detection

2026-04-02 · Weidong Tang, Hanbin Sun, Zihan Li, Yikai Wang 외 arxiv

Remote sensing change detection (CD) aims to identify where land-cover semantics change across time, but most existing methods still assume a fixed label space and therefore cannot answer arbitrary user-defined queries. …

Change Detection

RC-GeoCP: Geometric Consensus for Radar-Camera Collaborative Perception

2026-02-28 · Xiaokai Bai, Lianqing Zheng, Runwei Guan, Siyuan Cao 외 arxiv

Collaborative perception (CP) improves scene understanding through multi-agent information sharing, yet LiDAR-centric systems remain costly and vulnerable in adverse weather. Camera--4D radar offers a practical alternati…

Scene Understanding

From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs

2026-01-31 · Louis Schiekiera, Max Zimmer, Christophe Roux, Sebastian Pokutta 외 arxiv

We investigate the extent to which an LLM's hidden-state geometry can be recovered from its behavior in psycholinguistic experiments. Across eight instruction-tuned transformer models, we run two experimental paradigms -…

Consensus-Aware Visual-Semantic Embedding for Image-Text Matching

2020-07-17 · ECCV 2020 8 · Haoran Wang, Ying Zhang, Zhong Ji, Yanwei Pang 외

Image-text matching plays a central role in bridging vision and language. Most existing approaches only rely on the image-text instance pair to learn their representations, thereby exploiting their matching relationships…

Image CaptioningImage-text matchingRetrievalText Matching+1

Semantics-Aware Hierarchical Consensus Learning for Remote Sensing Image Classification

2025-10-06 · Giulio Weikmann, Gianmarco Perantoni, Lorenzo Bruzzone arxiv

Deep learning has become increasingly important in remote sensing image classification due to its ability to extract semantic information from complex data. Classification tasks often include predefined label hierarchies…

Remote Sensing Image ClassificationTime Series Classification