paper-with-me

Papers

Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D

2025-04-19 · Sergio Arnaud, Paul McVay, Ada Martin, Arjun Majumdar, Krishna Murthy Jatavallabhula, Phillip Thomas, Ruslan Partsey, Daniel Dugas, Abha Gejji, Alexander Sax, Vincent-Pierre Berges, Mikael Henaff, Ayush Jain, Ang Cao, Ishita Prasad, Mrinal Kalakrishnan, Michael Rabbat, Nicolas Ballas, Mido Assran, Oleksandr Maksymets, Aravind Rajeswaran, Franziska Meier

We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state-of-the-art on standard referential grounding benchmarks and showcases robust generalization capabilities. Notably, LOCATE 3D operates directly on sensor observation streams (posed RGB-D frames), enabling real-world deployment on robots and AR devices. Key to our approach is 3D-JEPA, a novel self-supervised learning (SSL) algorithm applicable to sensor point clouds. It takes as input a 3D pointcloud featurized using 2D foundation models (CLIP, DINO). Subsequently, masked prediction in latent space is employed as a pretext task to aid the self-supervised learning of contextualized pointcloud features. Once trained, the 3D-JEPA encoder is finetuned alongside a language-conditioned decoder to jointly predict 3D masks and bounding boxes. Additionally, we introduce LOCATE 3D DATASET, a new dataset for 3D referential grounding, spanning multiple capture setups with over 130K annotations. This enables a systematic study of generalization capabilities as well as a stronger model.

📄 PDF Abstract BibTeX arXiv:2504.14151

Code (1)

facebookresearch/locate-3d pytorch

Tasks

DecoderObject LocalizationSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Few-shot Object Localization

2024-03-19 · Yunhan Ren, Bo Li, Chengyang Zhang, Yong Zhang 외

Existing object localization methods are tailored to locate specific classes of objects, relying heavily on abundant labeled data for model optimization. However, acquiring large amounts of labeled data is challenging in…

Model OptimizationObjectObject CountingObject Localization

Peer-to-Peer Localization for Single-Antenna Devices

2020-12-10 · Xianan Zhang, Wei Wang, Xuedou Xiao, Hang Yang 외

Some important indoor localization applications, such as localizing a lost kid in a shopping mall, call for a new peer-to-peer localization technique that can localize an individual's smartphone or wearables by directly …

Indoor LocalizationTAG

MOGeo: Beyond One-to-One Cross-View Object Geo-localization

2026-03-14 · Bo Lv, Qingwang Zhang, Le Wu, Yuanyuan Li 외 arxiv

Cross-View Object Geo-Localization (CVOGL) aims to locate an object of interest in a query image within a corresponding satellite image. Existing methods typically assume that the query image contains only a single objec…

Visual and Object Geo-localization: A Comprehensive Survey

2021-12-30 · Daniel Wilson, Xiaohan Zhang, Waqas Sultani, Safwan Wshah

The concept of geo-localization refers to the process of determining where on earth some `entity' is located, typically using Global Positioning System (GPS) coordinates. The entity of interest may be an image, sequence …

3D Reconstructiongeo-localizationObjectSurvey

Object Manipulation via Visual Target Localization

2022-03-15 · Kiana Ehsani, Ali Farhadi, Aniruddha Kembhavi, Roozbeh Mottaghi

Object manipulation is a critical skill required for Embodied AI agents interacting with the world around them. Training agents to manipulate objects, poses many challenges. These include occlusion of the target object b…

Objectobject-detectionObject Detection