Agent-Centric Relation Graph for Object Visual Navigation
Object visual navigation aims to steer an agent toward a target object based on visual observations. It is highly desirable to reasonably perceive the environment and accurately control the agent. In the navigation task, we introduce an Agent-Centric Relation Graph (ACRG) for learning the visual representation based on the relationships in the environment. ACRG is a highly effective structure that consists of two relationships, i.e., the horizontal relationship among objects and the distance relationship between the agent and objects . On the one hand, we design the Object Horizontal Relationship Graph (OHRG) that stores the relative horizontal location among objects. On the other hand, we propose the Agent-Target Distance Relationship Graph (ATDRG) that enables the agent to perceive the distance between the target and objects. For ATDRG, we utilize image depth to obtain the target distance and imply the vertical location to capture the distance relationship among objects in the vertical direction. With the above graphs, the agent can perceive the environment and output navigation actions. Experimental results in the artificial environment AI2-THOR demonstrate that ACRG significantly outperforms other state-of-the-art methods in unseen testing environments.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectRelationVisual NavigationSimilar Papers 제목 키워드 기반
Zero-Shot Object Goal Visual Navigation With Class-Independent Relationship Network
This paper investigates the zero-shot object goal visual navigation problem. In the object goal visual navigation task, the agent needs to locate navigation targets from its egocentric visual input. "Zero-shot" means tha…
ObjectSemantic SimilaritySemantic Textual SimilarityVisual NavigationDeep Reinforcement Learning via Object-Centric Attention
Deep reinforcement learning agents, trained on raw pixel inputs, often fail to generalize beyond their training environments, relying on spurious correlations and irrelevant background details. To address this issue, obj…
Deep Reinforcement LearningInductive BiasObjectreinforcement-learning+1Image-Level Attentional Context Modeling Using Nested-Graph Neural Networks
We introduce a new scene graph generation method called image-level attentional context modeling (ILAC). Our model includes an attentional graph network that effectively propagates contextual information across the graph…
Graph GenerationGraph Neural NetworkObjectScene Graph GenerationUNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
Video Scene Graph Generation (VidSGG) aims to represent dynamic visual content by detecting objects and modeling their temporal interactions as structured graphs. Prior studies typically target either coarse-grained box-…
Video scene graph generationRepresentation LearningSEMBED: Semantic Embedding of Egocentric Action Videos
We present SEMBED, an approach for embedding an egocentric object interaction video in a semantic-visual graph to estimate the probability distribution over its potential semantic labels. When object interactions are ann…
General ClassificationObject