paper-with-me

홈 › Papers

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization

2026-06-02 · Tewodros Ayalew, Matthew Jeung, Samuel Wheeler, Xiao Zhang, Andre de la Cruz Arce, Kaylene Stocking, Michael Maire, Matthew R. Walter arxiv

We introduce CLAW, a fully end-to-end self-supervised framework for learning a world model jointly with continuous latent action representations directly from action-free videos. Our approach leverages adversarial latent regularization and diffusion-based video generation to capture structured and semantically meaningful action representations while modeling rich, predictive environment dynamics, without relying on any action labels or annotations. By simultaneously training the Latent Action Model and world model, CLAW learns to reason about how inferred actions induce environment transitions from visual observations alone. We show that the resulting latent action world model supports both imitation learning from observation and goal-directed planning. In imitation learning, latent actions extracted from raw videos enable behavior cloning. For planning, CLAW generates sequences of latent actions and maps them to executable actions to reach desired goals. Extensive experiments across diverse tasks and embodiments demonstrate that CLAW produces semantically meaningful latent action representations, supports effective action transfer, and enables planning and imitation from observation, outperforming existing methods.

📄 PDF Abstract BibTeX arXiv:2606.04130

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VisionClaw: Always-On AI Agents through Smart Glasses

2026-04-03 · Xiaoan Liu, DaeHo Lee, Eric J Gonzalez, Mar Gonzalez-Franco 외 arxiv

We present VisionClaw, an always-on wearable AI agent that integrates live egocentric perception with agentic task execution. Running on Meta Ray-Ban smart glasses, VisionClaw continuously perceives real-world context an…

CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping

2025-09-17 · Zijian An, Ran Yang, Yiming Feng, Lifeng Zhou arxiv

Vision-language-action (VLA) models have recently emerged as a promising paradigm for robotic control, enabling end-to-end policies that ground natural language instructions into visuomotor actions. However, current VLAs…

Robotic Grasping

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

2026-07-14 · Yunxin Li, Jinchao Li, Shibo Su, Zhenran Xu 외 arxiv

OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation …

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

2026-04-09 · Ziyu Ma, Shidong Yang, Yuxiang Ji, Xucong Wang 외 arxiv

Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar workflows, tool usage patterns, and failu…

ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution

2026-08-06 · Jiacheng Wei, Zhaoxin Fan, Xin Wen, Yuqin Lan 외 arxiv

General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments. On-chain execution is stateful, adversarial, and e…