paper-with-me

Papers

DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships

2024-10-14 · Zhang Wan, Sheng Tang, Jiawei Wei, Ruize Zhang, Juan Cao

In recent years, diffusion models have achieved tremendous success in the field of video generation, with controllable video generation receiving significant attention. However, existing control methods still face two limitations: Firstly, control conditions (such as depth maps, 3D Mesh) are difficult for ordinary users to obtain directly. Secondly, it's challenging to drive multiple objects through complex motions with multiple trajectories simultaneously. In this paper, we introduce DragEntity, a video generation model that utilizes entity representation for controlling the motion of multiple objects. Compared to previous methods, DragEntity offers two main advantages: 1) Our method is more user-friendly for interaction because it allows users to drag entities within the image rather than individual pixels. 2) We use entity representation to represent any object in the image, and multiple objects can maintain relative spatial relationships. Therefore, we allow multiple trajectories to control multiple objects in the image with different levels of complexity simultaneously. Our experiments validate the effectiveness of DragEntity, demonstrating its excellent performance in fine-grained control in video generation.

📄 PDF Abstract BibTeX arXiv:2410.10751

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

2025-07-08 · Zhenghao Zhang, Junchao Liao, Xiangyu Meng, Long Qin 외

Recent advances in diffusion transformer models for motion-guided video generation, such as Tora, have shown significant progress. In this paper, we present Tora2, an enhanced version of Tora, which introduces several de…

Video Generation

The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text

2025-12-18 · Hanlin Wang, Hao Ouyang, Qiuyu Wang, Yue Yu 외 arxiv

We present WorldCanvas, a framework for promptable world events that enables rich, user-directed simulation by combining text, trajectories, and reference images. Unlike text-only approaches and existing trajectory-contr…

Visual Grounding

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

2026-04-10 · Junchao Liao, Zhenghao Zhang, Xiangyu Meng, Litao Li 외 arxiv

Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible motion-sound relations remains challenging. Existing methods often produ…

Audio GenerationVideo Generation

Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training

2025-09-08 · Ruicheng Zhang, Jun Zhou, Zunnan Xu, Zihao Liu 외 arxiv

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated …

Video Generation

TrajectoryMover: Generative Movement of Object Trajectories in Videos

2026-03-31 · Kiran Chhatre, Hyeonho Jeong, Yulia Gryaditskaya, Christopher E. Peters 외 arxiv

Generative video editing has enabled several intuitive editing operations for short video clips that would previously have been difficult to achieve, especially for non-expert editors. Existing methods focus on prescribi…