paper-with-me

Papers

WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language

2023-04-12 · Zhenxiang Lin, Xidong Peng, Peishan Cong, Ge Zheng, Yujin Sun, Yuenan Hou, Xinge Zhu, Sibei Yang, Yuexin Ma

We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a novel method, dubbed WildRefer, for this task by fully utilizing the rich appearance information in images, the position and geometric clues in point cloud as well as the semantic knowledge of language descriptions. Besides, we propose two novel datasets, i.e., STRefer and LifeRefer, which focus on large-scale human-centric daily-life scenarios accompanied with abundant 3D object and natural language annotations. Our datasets are significant for the research of 3D visual grounding in the wild and has huge potential to boost the development of autonomous driving and service robots. Extensive experiments and ablation studies demonstrate that our method achieves state-of-the-art performance on the proposed benchmarks. The code is provided in https://github.com/4DVLab/WildRefer.

📄 PDF Abstract BibTeX arXiv:2304.05645

Code (1)

4dvlab/wildrefer 공식 구현 pytorch

Tasks

3D visual groundingAutonomous DrivingObject LocalizationVisual Grounding

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

RDNet: Region Proportion-Aware Dynamic Adaptive Salient Object Detection Network in Optical Remote Sensing Images

2026-03-12 · Bin Wan, Runmin Cong, Xiaofei Zhou, Hao Fang 외 arxiv

Salient object detection (SOD) in remote sensing images faces significant challenges due to large variations in object sizes, the computational cost of self-attention mechanisms, and the limitations of CNN-based extracto…

Salient Object DetectionObject Localization

Multi-scale Multi-instance Visual Sound Localization and Segmentation

2024-08-31 · Shentong Mo, Haofan Wang

Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global …

Object Localization

Multi-scale Volumes for Deep Object Detection and Localization

2015-05-14 · Eshed Ohn-Bar, M. M. Trivedi

This study aims to analyze the benefits of improved multi-scale reasoning for object detection and localization with deep convolutional neural networks. To that end, an efficient and general object detection framework wh…

Objectobject-detectionObject Detection

Lightweight Object-level Topological Semantic Mapping and Long-term Global Localization based on Graph Matching

2022-01-16 · Fan Wang, Chaofan Zhang, Fulin Tang, Hongkui Jiang 외

Mapping and localization are two essential tasks for mobile robots in real-world applications. However, largescale and dynamic scenes challenge the accuracy and robustness of most current mature solutions. This situation…

Graph MatchingManagement

Semantic-LiDAR-Inertial-Wheel Odometry Fusion for Robust Localization in Large-Scale Dynamic Environments

2025-09-18 · Haoxuan Jiang, Peicong Qian, Yusen Xie, Linwei Zheng 외 arxiv

Reliable, drift-free global localization presents significant challenges yet remains crucial for autonomous navigation in large-scale dynamic environments. In this paper, we introduce a tightly-coupled Semantic-LiDAR-Ine…