paper-with-me

Papers

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory

2026-06-16 · Zhexiao Xiong, Yizhi Song, Hao Kang, Qing Yan, Liming Jiang, Jenson Yang, Zhoujie Fu, Stathi Fotiadis, Angtian Wang, Zichuan Liu, Bo Liu, Yiding Yang, Xin Lu, Nathan Jacobs arxiv

Interactive world models aim to simulate environment dynamics under real-time user actions. However, their action vocabulary is largely confined to navigation: most actions correspond to motion (e.g., walk, turn, look around), while interaction with objects in the scene (e.g., pick up plates, open doors, or trigger physical responses) is either absent, restricted to game domains, or relegated to prompt-to-full-video scenarios. The resulting worlds are visually explorable but not truly actionable. In this work, we present ActWorld, an interactive world model that extends prior navigation-centric generators to support mid-rollout object interaction within a chunk-autoregressive framework. We argue that the navigation-interaction gap stems from two bottlenecks. First, a data bottleneck: the lack of human-object interaction data with accurate, dense labels. Second, a memory bottleneck: recency-biased history compression in existing world models discards the event-transition frames that causally determine subsequent object states, leading to an action-forgetting pathology. On the data side, we construct a 100K interaction video dataset, each annotated with per-chunk captions via chain-of-thought reasoning. On the model side, we introduce a hierarchical action-aware memory design that routes history compression by interaction importance, complemented by a persistent memory bank that maintains event-update and object-identity tokens across long rollouts. Experiments show that ActWorld supports both flexible navigation and rich object interaction within a single model, substantially improving interaction fidelity over navigation-only baselines without sacrificing viewpoint control. Project page is available at https://interactwm.github.io/ActWorld.

📄 PDF Abstract BibTeX arXiv:2606.17730

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

2025-07-29 · HunyuanWorld Team, Zhenwei Wang, Yuhao Liu, Junta Wu 외 arxiv

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods…

WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes

2026-05-15 · Jichen Hu, Jiawei Guo, Jiazhong Cen, Chen Yang 외 arxiv

Recent 3D world modeling systems based on generative scene synthesis, such as Marble, can create coherent and explorable 3D environments, yet their outputs are typically static monolithic assets with limited editability …

ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models

2026-04-14 · Xinliang Wang, Yifeng Shi, Zhenyu Wu arxiv

3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constraints. Current generative restoration approaches are often limited b…

Novel View Synthesis3D ReconstructionVideo Generation

Augmented Math: Authoring AR-Based Explorable Explanations by Augmenting Static Math Textbooks

2023-07-30 · Neil Chulpongsatorn, Mille Skovhus Lunding, Nishan Soni, Ryo Suzuki

We introduce Augmented Math, a machine learning-based approach to authoring AR explorable explanations by augmenting static math textbooks without programming. To augment a static document, our system first extracts math…

MathOptical Character RecognitionOptical Character Recognition (OCR)

ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation

2026-06-11 · Zhiyuan Zhang, Pokuang Zhou, Kaidi Zhang, Adeesh Desai 외 arxiv

Contact-rich manipulation requires world models to reason over complex contact dynamics from multimodal sensory observations. However, it remains unclear which representation properties fundamentally support stable long-…