paper-with-me

Papers

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

2026-03-16 · Jinguang Tong, Jinbo Wu, Kaisiyuan Wang, Zhelun Shen, Xuan Huang, Mochu Xiang, Xuesong Li, Yingying Li, Haocheng Feng, Chen Zhao, Hang Zhou, Wei He, Chuong Nguyen, Jingdong Wang, Hongdong Li arxiv

Human-Object Interaction (HOI) video reenactment aims to transfer the interaction dynamics of a source video to a novel target object while preserving realistic hand-object coordination. Existing methods typically rely on sparse 2D motion controls and monocular references, which are insufficient for complex out-of-plane motion and large viewpoint changes. We present MVHOI, a two-stage framework combining implicit motion extraction, 3D-aware multi-view reasoning, and video generation. In the first stage, a motion extractor encodes object dynamics into implicit motion descriptors. Conditioned on these descriptors, our Motion-Driven Object Prior (MDOP) module queries a 3D foundation model over multi-view references of the target object and autoregressively predicts coarse object anchors, a sequence of images that track the object's evolving orientation and appearance under the source motion without any explicit pose estimation. In the second stage, a DiT-based video generation model uses these anchors as structural guidance and the multi-view references as appearance guidance. We further reuse cross-view attention from MDOP as a soft attention bias to reduce reference-view confusion. For long videos, a cross-iterative inference strategy refreshes subsequent object priors using refined video outputs. Experiments demonstrate consistent improvements over state-of-the-art methods in object fidelity, motion consistency, visual quality, and interaction realism.

📄 PDF Abstract BibTeX arXiv:2603.14686

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Dynamic Black-Litterman

2024-04-29 · Anas Abdelhakmi, Andrew Lim

The Black-Litterman model is a framework for incorporating forward-looking expert views in a portfolio optimization problem. Existing work focuses almost exclusively on single-period problems with the forecast horizon ma…

modelPortfolio Optimization

MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models

2025-12-03 · Shaoheng Fang, Chaohui Yu, Fan Wang, Qixing Huang arxiv

We introduce MVRoom, a controllable novel view synthesis (NVS) pipeline for 3D indoor scenes that uses multi-view diffusion conditioned on a coarse 3D layout. MVRoom employs a two-stage design in which the 3D layout is u…

Novel View SynthesisScene Generation

BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks

2026-02-03 · Yixiang Chen, Peiyan Li, Jiabing Yang, Keji He 외 arxiv

Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual and motion priors. However, they still fac…

Video Generation

A Unified Framework for Diffusion Bridge Problems: Flow Matching and Schrödinger Matching into One

2025-03-27 · Minyoung Kim

The bridge problem is to find an SDE (or sometimes an ODE) that bridges two given distributions. The application areas of the bridge problem are enormous, among which the recent generative modeling (e.g., conditional or …

Image GenerationUnconditional Image Generation

Generative Adversarial Frontal View to Bird View Synthesis

2018-08-01 · Xinge Zhu, Zhichao Yin, Jianping Shi, Hongsheng Li 외

Environment perception is an important task with great practical value and bird view is an essential part for creating panoramas of surrounding environment. Due to the large gap and severe deformation between the frontal…

Bird View SynthesisHomography EstimationTranslation