paper-with-me

홈 › Papers

ContactFlow: A video action conditioning that transfers across embodiments

2026-07-29 · Sami Azirar, Enrico Pallotta, Jan Nogga, Jürgen Gall, Sven Behnke, Hermann Blum arxiv

World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact. Further, their action conditioning is often constrained to specific embodiments such as parallel grippers. We propose \emph{Contact Flow}, an embodiment-agnostic action representation that encodes manipulation through the trajectory of 3D contact points between an actor and a target object. By discarding actor-specific appearance and kinematics, Contact Flow provides a shared conditioning signal for both human demonstrations and robotic execution. Therefore, we can train a large-scale video generative model on both human and robotic object interaction videos conditioned on Contact Flow, yielding a world model that predicts physically plausible manipulation outcomes. We integrate this model into a propose-imagine-verify-act pipeline, where generated rollouts are assessed by a vision-language model before execution. Experiments on the DROID dataset and real-world tabletop manipulation tasks demonstrate that Contact Flow enables transfer between human demonstrations and different robotic embodiments.

📄 PDF Abstract BibTeX arXiv:2607.26579

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation

2026-05-23 · Ofir Abramovich, Nadav Z. Cohen, Adi Rosenthal, Ariel Shamir arxiv

Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs. However, existing conditioning methods often require additional trai…

Video Generation

ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation

2026-05-07 · Omar El Khalifi, Thomas Rossi, Oscar Fossey, Thibault Fouque 외 arxiv

For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video gene…

Video Generation

VISTA: Enhancing Visual Conditioning via Track-Following Preference Optimization in Vision-Language-Action Models

2026-02-04 · Yiye Chen, Yanan Jian, Xiaoyi Dong, Shuxin Cao 외 arxiv

Vision-Language-Action (VLA) models have demonstrated strong performance across a wide range of robotic manipulation tasks. Despite the success, extending large pretrained Vision-Language Models (VLMs) to the action spac…

Utonia: Toward One Encoder for All Point Clouds

2026-03-03 · Yujia Zhang, Xiaoyang Wu, Yunhan Yang, Xianzhe Fan 외 arxiv

We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self-supervised point …

Multimodal ReasoningAutonomous DrivingSpatial ReasoningPoint Clouds

Aurora: Unified Video Editing with a Tool-Using Agent

2026-05-18 · Yongsheng Yu, Ziyun Zeng, Zhiyuan Xiao, Zhenghong Zhou 외 arxiv

Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and reference images, and one set of weights covers replacement, removal, s…

Style Transfer