paper-with-me

홈 › Papers

Point-to-Point: Sparse Motion Guidance for Controllable Video Editing

2025-11-23 · Yeji Song, Jaehyun Lee, Mijin Koo, JunHoo Lee, Nojun Kwak arxiv

Accurately preserving motion while editing a subject remains a core challenge in video editing tasks. Existing methods often face a trade-off between edit and motion fidelity, as they rely on motion representations that are either overfitted to the layout or only implicitly defined. To overcome this limitation, we revisit point-based motion representation. However, identifying meaningful points remains challenging without human input, especially across diverse video scenarios. To address this, we propose a novel motion representation, anchor tokens, that capture the most essential motion patterns by leveraging the rich prior of a video diffusion model. Anchor tokens encode video dynamics compactly through a small number of informative point trajectories and can be flexibly relocated to align with new subjects. This allows our method, Point-to-Point, to generalize across diverse scenarios. Extensive experiments demonstrate that anchor tokens lead to more controllable and semantically aligned video edits, achieving superior performance in terms of edit and motion fidelity.

📄 PDF Abstract BibTeX arXiv:2511.18277

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DisPose: Disentangling Pose Guidance for Controllable Human Image Animation

2024-12-12 · Hongxiang Li, Yaowei Li, Yuhang Yang, Junjie Cao 외

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to …

Image Animation

Controllable Human-Object Interaction Synthesis

2023-12-06 · Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu 외

Synthesizing semantic-aware, long-horizon, human-object interaction is critical to simulate realistic human behaviors. In this work, we address the challenging problem of generating synchronized object motion and human m…

Human-Object Interaction DetectionObject

Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D Generation

2023-07-26 · Chaohui Yu, Qiang Zhou, Jingliang Li, Zhe Zhang 외

Text-to-3D generation has recently garnered significant attention, fueled by 2D diffusion models trained on billions of image-text pairs. Existing methods primarily rely on score distillation to leverage the 2D diffusion…

3D GenerationNeRFText to 3D

Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling

2024-01-29 · Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian 외

We introduce Motion-I2V, a novel framework for consistent and controllable image-to-video generation (I2V). In contrast to previous methods that directly learn the complicated image-to-video mapping, Motion-I2V factorize…

Image to Video GenerationVideo Generation

Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text

2026-06-03 · Jaeyeong Kim, Ines Kim, Jahyeok Koo, Seungryong Kim arxiv

We introduce T2Mo, a feed-forward framework for controllable dynamic 3D shape generation conditioned on 3D trajectories and text. Due to the inherent ambiguity of language, generating precisely intended motions using tex…

Video Generation