paper-with-me

홈 › Papers

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding

2025-12-28 · Wenyuan Huang, Zhenyu Zhang, Zhao Wang, Zhou Wei, Ting Huang, Fang Zhao, Jian Yang arxiv

3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing supervised methods are limited by generalization and recent zero-shot methods typically rely on a predefined Object Lookup Table (OLT) to query Visual Language Models (VLMs) for reasoning about object locations via a single step grounding, which limits the applications in scenarios with undefined targets and complex queries. To address these problems, we present OpenGround, a novel zero-shot framework for open-world 3D visual grounding that remains compatible with recent zero-shot methods. OpenGround integrates Task-Chain Planning to decompose a query into a plan of context-to-target sub-goals for progressive grounding, and Context-Guided Perception to perceive novel objects online under context guidance from the task chain. We also propose a new dataset named OpenTarget, which contains over 7000 object-description pairs to mimic open-world evaluation. Extensive experiments demonstrate that OpenGround achieves competitive performance on Nr3D, state-of-the-art on ScanRefer, and delivers a substantial 17.6\% improvement on OpenTarget. Project Page at https://why-102.github.io/openground.io/.

📄 PDF Abstract BibTeX arXiv:2512.23020

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Perception-and-Energy-aware Motion Planning for UAV using Learning-based Model under Heteroscedastic Uncertainty

2023-09-25 · Reiya Takemura, Genya Ishigami

Global navigation satellite systems (GNSS) denied environments/conditions require unmanned aerial vehicles (UAVs) to energy-efficiently and reliably fly. To this end, this study presents perception-and-energy-aware motio…

Motion PlanningTrajectory Planning

UItron: Foundational GUI Agent with Advanced Perception and Planning

2025-08-29 · Zhixiong Zeng, Jing Huang, Liming Zheng, Wenkang Han 외 arxiv

GUI agent aims to enable automated operations on Mobile/PC devices, which is an important task toward achieving artificial general intelligence. The rapid advancement of VLMs accelerates the development of GUI agents, ow…

Reinforcement Learning

G-MAPP: GPU-accelerated Multi-Agent Planning and Perception for Reactive Motion Generation

2026-06-10 · Tanmay Bishnoi, Riddhiman Laha, Tobias Löw, Jose Alex Chandy 외 arxiv

Reactive motion generation in unstructured environments remains an open challenge in robotics. Due to the computational complexity of collision-free motion generation, existing methods either generate global trajectories…

Trajectory Planning

World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

2025-07-01 · Yupeng Zheng, Pengxuan Yang, Zebin Xing, Qichao Zhang 외

End-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information. A critical research challenge arises: co…

Autonomous DrivingNavSimSelf-Supervised Learning

V2X-ReaLO: An Open Online Framework and Dataset for Cooperative Perception in Reality

2025-03-13 · Hao Xiang, Zhaoliang Zheng, Xin Xia, Seth Z. Zhao 외

Cooperative perception enabled by Vehicle-to-Everything (V2X) communication holds significant promise for enhancing the perception capabilities of autonomous vehicles, allowing them to overcome occlusions and extend thei…

Autonomous Vehicles