paper-with-me

Papers

Objects Matter: Learning Object Relation Graph for Robust Camera Relocalization

2022-05-26 · Chengyu Qiao, Zhiyu Xiang, Xinglu Wang

Visual relocalization aims to estimate the pose of a camera from one or more images. In recent years deep learning based pose regression methods have attracted many attentions. They feature predicting the absolute poses without relying on any prior built maps or stored images, making the relocalization very efficient. However, robust relocalization under environments with complex appearance changes and real dynamics remains very challenging. In this paper, we propose to enhance the distinctiveness of the image features by extracting the deep relationship among objects. In particular, we extract objects in the image and construct a deep object relation graph (ORG) to incorporate the semantic connections and relative spatial clues of the objects. We integrate our ORG module into several popular pose regression models. Extensive experiments on various public indoor and outdoor datasets demonstrate that our method improves the performance significantly and outperforms the previous approaches.

📄 PDF Abstract BibTeX arXiv:2205.13280

Code (0)

등록된 구현이 없습니다.

Tasks

Camera RelocalizationregressionRelation

Similar Papers 제목 키워드 기반

Incremental Real-Time Multibody VSLAM with Trajectory Optimization Using Stereo Camera

2016-08-02 · N. Dinesh Reddy, Iman Abbasnejad, Sheetal Reddy, Amit Kumar Mondal 외

Real time outdoor navigation in highly dynamic environments is an crucial problem. The recent literature on real time static SLAM don't scale up to dynamic outdoor environments. Most of these methods assume moving object…

Motion Segmentation

MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known Objects

2026-03-24 · Shiyu Li, Hannah Schieber, Kristoffer Waldow, Benjamin Busam 외 arxiv

Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achieved by combining contextual information…

Camera Pose Estimation

EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories

2026-06-06 · Hassan Jaber, Refinath S N, Luca Cagliero, Christopher E. Mower 외 arxiv

Spatial grounding remains a key limitation of vision-language-action (VLA) systems for robotic manipulation. While current models can recognize objects and follow language instructions, they often lack an explicit repres…

Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos

2026-03-13 · Rohith Peddi, Saurabh, Shravan Shanmugam, Likhitha Pallapothula 외 arxiv

Spatio-temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame-centric: they reason only about currently visible objects, disca…

Scene Graph GenerationScene Understanding3D Reconstruction

Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph

2025-07-16 · Sergey Linok, Gleb Naumov arxiv

We propose OVIGo-3DHSG method - Open-Vocabulary Indoor Grounding of objects using 3D Hierarchical Scene Graph. OVIGo-3DHSG represents an extensive indoor environment over a Hierarchical Scene Graph derived from sequences…

Spatial Reasoning