paper-with-me

홈 › Papers

AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors

2024-03-15 · Yucen Wang, Shenghua Wan, Le Gan, Shuai Feng, De-Chuan Zhan

Model-based methods have significantly contributed to distinguishing task-irrelevant distractors for visual control. However, prior research has primarily focused on heterogeneous distractors like noisy background videos, leaving homogeneous distractors that closely resemble controllable agents largely unexplored, which poses significant challenges to existing methods. To tackle this problem, we propose Implicit Action Generator (IAG) to learn the implicit actions of visual distractors, and present a new algorithm named implicit Action-informed Diverse visual Distractors Distinguisher (AD3), that leverages the action inferred by IAG to train separated world models. Implicit actions effectively capture the behavior of background distractors, aiding in distinguishing the task-irrelevant components, and the agent can optimize the policy within the task-relevant state space. Our method achieves superior performance on various visual control tasks featuring both heterogeneous and homogeneous distractors. The indispensable role of implicit actions learned by IAG is also empirically validated.

📄 PDF Abstract BibTeX arXiv:2403.09976

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Implicit Behavioral Cloning

2021-09-01 · Pete Florence, Corey Lynch, Andy Zeng, Oscar Ramirez 외

We find that across a wide range of robot policy learning scenarios, treating supervised policy learning with an implicit model generally performs better, on average, than commonly used explicit models. We present extens…

D4RL

VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

2022-09-30 · Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani 외

Reward and representation learning are two long-standing challenges for learning an expanding set of robot manipulation skills from sensory observations. Given the inherent cost and scarcity of in-domain, task-specific r…

Offline RLOpen-Ended Question AnsweringRepresentation LearningRobot Manipulation

Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

2025-07-01 · Tao Lin, Gen Li, Yilei Zhong, Yanwen Zou 외 arxiv

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-L…

Depth EstimationPoint Clouds

From Prediction to Self: Developmental Conditions for Agency in Minimal Neural Systems

2026-06-04 · Evan Ye arxiv

How does a system that merely predicts the world come to distinguish its own causal influence from everything else? We trace this transition in a minimal 192-dimensional GRU through a developmental sequence -- 6 experime…

UniAR: A Unified model for predicting human Attention and Responses on visual content

2023-12-15 · Peizhao Li, Junfeng He, Gang Li, Rachit Bhargava 외

Progress in human behavior modeling involves understanding both implicit, early-stage perceptual behavior, such as human attention, and explicit, later-stage behavior, such as subjective preferences or likes. Yet most pr…