paper-with-me

홈 › Papers

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

2025-10-11 · Feng Xiao, Hongbin Xu, Hai Ci, Wenxiong Kang arxiv

Localizing 3D objects using natural language is essential for robotic scene understanding. The descriptions often involve multiple spatial relationships to distinguish similar objects, making 3D-language alignment difficult. Current methods only model relationships for pairwise objects, ignoring the global perceptual significance of n-ary combinations in multi-modal relational understanding. To address this, we propose a novel progressive relational learning framework for 3D object grounding. We extend relational learning from binary to n-ary to identify visual relations that match the referential description globally. Given the absence of specific annotations for referred objects in the training data, we design a grouped supervision loss to facilitate n-ary relational learning. In the scene graph created with n-ary relationships, we use a multi-modal network with hybrid attention mechanisms to further localize the target within the n-ary combinations. Experiments and ablation studies on the ReferIt3D and ScanRefer benchmarks demonstrate that our method outperforms the state-of-the-art, and proves the advantages of the n-ary relational perception in 3D localization.

📄 PDF Abstract BibTeX arXiv:2510.10194

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

Progressive Prompt-Guided Cross-Modal Reasoning for Referring Image Segmentation

2026-03-30 · Jiachen Li, Hongyun Wang, Jinyu Xu, Wenbo Jiang 외 arxiv

Referring image segmentation aims to localize and segment a target object in an image based on a free-form referring expression. The core challenge lies in effectively bridging linguistic descriptions with object-level v…

Semantic SegmentationInstance SegmentationReferring ExpressionImage Segmentation

Progressive Local Alignment for Medical Multimodal Pre-training

2025-02-25 · Huimin Yan, Xian Yang, Liang Bai, Jiye Liang

Local alignment between medical images and text is essential for accurate diagnosis, though it remains challenging due to the absence of natural local pairings and the limitations of rigid region recognition methods. Tra…

Contrastive LearningImage-text Retrievalobject-detectionObject Detection+4

MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning

2025-09-26 · Lihao Zheng, Jiawei Chen, Xintian Shen, Hao Ma 외 arxiv

Multi-image reasoning and grounding require understanding complex cross-image relationships at both object levels and image levels. Current Large Visual Language Models (LVLMs) face two critical challenges: the lack of c…

Reinforcement Learning

Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models

2026-04-07 · Sakshi Agarwal, Aishik Konwer, Ankit Parag Shah arxiv

Video Anomaly Detection (VAD) has traditionally been framed as binary classification or outlier detection, providing neither interpretable reasoning nor precise spatial localization of anomalous events. While Vision-Lang…

Video Anomaly DetectionAnomaly ClassificationDomain GeneralizationBinary Classification

3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection

2022-04-13 · CVPR 2022 1 · Junyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao 외

3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage paradigm, i.e., language-irrelevant detecti…

3D visual groundingVisual Grounding