paper-with-me

Papers

Environment Predictive Coding for Embodied Agents

2021-02-03 · Santhosh K. Ramakrishnan, Tushar Nagarajan, Ziad Al-Halah, Kristen Grauman

We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for images, we aim to jointly encode a series of images gathered by an agent as it moves about in 3D environments. We learn these representations via a zone prediction task, where we intelligently mask out portions of an agent's trajectory and predict them from the unmasked portions, conditioned on the agent's camera poses. By learning such representations on a collection of videos, we demonstrate successful transfer to multiple downstream navigation-oriented tasks. Our experiments on the photorealistic 3D environments of Gibson and Matterport3D show that our method outperforms the state-of-the-art on challenging tasks with only a limited budget of experience.

📄 PDF Abstract BibTeX arXiv:2102.02337

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Environment Predictive Coding for Visual Navigation

2021-09-29 · ICLR 2022 4 · Santhosh Kumar Ramakrishnan, Tushar Nagarajan, Ziad Al-Halah, Kristen Grauman

We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for individual images, we aim t…

Representation LearningSelf-Supervised LearningVisual Navigation

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning

2026-05-10 · Haoqiang Kang, Xiaokang Ye, Yuhan Liu, Siddhant Hitesh Mantri 외 arxiv

LLM/VLM-based digital agents have advanced rapidly thanks to scalable sandboxes for coding, web navigation, and computer use, which provide rich interactive training grounds. In contrast, embodied agents still lack abund…

3D Generation

3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding

2026-04-09 · Makanjuola Ogunleye, Eman Abdelrahman, Ismini Lourentzou arxiv

Large multimodal models are increasingly used as the reasoning core of embodied agents operating in 3D environments, yet they remain prone to hallucinations that can produce unsafe and ungrounded decisions. Existing infe…

MaskViT: Masked Visual Pre-Training for Video Prediction

2022-06-23 · Agrim Gupta, Stephen Tian, Yunzhi Zhang, Jiajun Wu 외

The ability to predict future visual observations conditioned on past observations and motor commands can enable embodied agents to plan solutions to a variety of tasks in complex environments. This work shows that we ca…

PredictionSchedulingVideo Prediction

Autonomous Embodied Agents: When Robotics Meets Deep Learning Reasoning

2025-05-02 · Roberto Bigazzi

The increase in available computing power and the Deep Learning revolution have allowed the exploration of new topics and frontiers in Artificial Intelligence research. A new field called Embodied Artificial Intelligence…

Deep Learning