paper-with-me

홈 › Papers

DC-WAM: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models

2026-07-28 · Haoyuan Ji, Lingxiang Fan, Shang Su, Yinqiao Lu, Mengkai Shi, Jun Gao, Shuo Feng arxiv

World-Action Models (WAMs) augment robot policies with future visual prediction, but it remains unclear what the visual modality should learn for control. While photorealistic future prediction provides dense supervision, it also incurs substantial computation and can allocate capacity to texture, illumination, and background variations that are only weakly related to action selection. Recent efficient WAM variants suggest that the main benefit of the video branch may not lie in the rendered future itself, but in the control-relevant visual representations induced during training. In this work, we revisit future video prediction from a dynamic-centric perspective and ask whether an existing RGB-based WAM can be redirected from appearance-dominated reconstruction toward interaction-induced visual dynamics without introducing additional modality-specific predictions or online inputs at deployment. We propose DC-WAM, a dynamic-centric WAM framework that redistributes supervision and computation in the RGB video branch. At the supervision level, DC-WAM combines temporal-difference flow matching with trajectory-guided weighting, emphasizing dense temporal changes and localized regions where the gripper, manipulated objects, and contact areas move. At the reasoning level, DynaRoute predicts token-wise dynamic relevance and converts it into an attention bias, guiding the model toward control-relevant future tokens. Experiments in simulation and on real-world manipulation tasks show that DC-WAM consistently improves policy performance, especially under out-of-distribution perturbations in lighting, object appearance, and background texture.

📄 PDF Abstract BibTeX arXiv:2607.25918

Code (0)

등록된 구현이 없습니다.

Tasks

Video Prediction

Similar Papers 제목 키워드 기반

Weakly Supervised Concept Learning for Object-centric Visual Reasoning

2026-05-05 · Sparsh Tiwari, Bettina Finzel, Gesina Schwalbe arxiv

Neurosymbolic systems promise to combine deep neural network's (DNN) processing of raw sensor inputs with few-shot performance of symbolic artificial intelligence. Two-stage approaches explicitly decouple DNN based perce…

Inductive logic programmingDomain GeneralizationVisual Reasoning

SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos

2022-06-15 · Gamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff 외

The visual world can be parsimoniously characterized in terms of distinct entities with sparse interactions. Discovering this compositional structure in dynamic visual scenes has proven challenging for end-to-end compute…

ObjectSemantic Segmentation

World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

2026-04-29 · Wanyue Zhang, Wenxiang Wu, Wang Xu, Jiaxin Luo 외 arxiv

Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent…

Spatial Reasoning

nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving

2026-05-29 · Zhiyu Huang, Johnson Liu, Rui Song, Zewei Zhou 외 arxiv

Reasoning is essential for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relations, infer agent interactions, and make safe decisions. However, existi…

Visual Question AnsweringAutonomous DrivingSpatial Reasoning

HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision

2025-07-02 · Shengli Zhou, Jianuo Zhu, Qilin Huang, Fangjing Wang 외 arxiv

3D Visual Question-Answering (3D VQA) is pivotal for models to perceive the physical world and perform spatial reasoning. Answer-centric supervision is a commonly used training method for 3D VQA models. Many models that …

Spatial Reasoning