paper-with-me

홈 › Papers

Target-Aware Video Diffusion Models

2025-03-24 · Taeksoo Kim, Hanbyul Joo

We present a target-aware video diffusion model that generates videos from an input image in which an actor interacts with a specified target while performing a desired action. The target is defined by a segmentation mask and the desired action is described via a text prompt. Unlike existing controllable image-to-video diffusion models that often rely on dense structural or motion cues to guide the actor's movements toward the target, our target-aware model requires only a simple mask to indicate the target, leveraging the generalization capabilities of pretrained models to produce plausible actions. This makes our method particularly effective for human-object interaction (HOI) scenarios, where providing precise action guidance is challenging, and further enables the use of video diffusion models for high-level action planning in applications such as robotics. We build our target-aware model by extending a baseline model to incorporate the target mask as an additional input. To enforce target awareness, we introduce a special token that encodes the target's spatial information within the text prompt. We then fine-tune the model with our curated dataset using a novel cross-attention loss that aligns the cross-attention maps associated with this token with the input target mask. To further improve performance, we selectively apply this loss to the most semantically relevant transformer blocks and attention regions. Experimental results show that our target-aware model outperforms existing solutions in generating videos where actors interact accurately with the specified targets. We further demonstrate its efficacy in two downstream applications: video content creation and zero-shot 3D HOI motion synthesis.

📄 PDF Abstract BibTeX arXiv:2503.18950

Code (0)

등록된 구현이 없습니다.

Tasks

Human-Object Interaction DetectionMotion Synthesis

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

GenVideo: One-shot Target-image and Shape Aware Video Editing using T2I Diffusion Models

2024-04-18 · Sai Sree Harsha, Ambareesh Revanur, Dhwanit Agarwal, Shradha Agrawal

Video editing methods based on diffusion models that rely solely on a text prompt for the edit are hindered by the limited expressive power of text prompts. Thus, incorporating a reference target image as a visual guide …

Video Editing

PE-Field 4D: Video Generation Models as Canvas

2026-07-17 · Yunpeng Bai, Haoxiang Li, Qixing Huang arxiv

Diffusion Transformers have recently achieved strong performance in video generation, yet controlling scene geometry under viewpoint changes and camera motion remains challenging. In this work, we revisit the role of pos…

Video Generation

MotionAdapter: Video Motion Transfer via Content-Aware Attention Customization

2026-01-05 · Zhexin Zhang, Yangyang Xu, Yifeng Zhu, Long Chen 외 arxiv

Recent advances in diffusion-based text-to-video models, particularly those built on the diffusion transformer architecture, have achieved remarkable progress in generating high-quality and temporally coherent videos. Ho…

Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning

2026-05-20 · Zhangchi Hu, Wenzhang Sun, Xiangchen Yin, Jiahui Yuan 외 arxiv

Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions while synthesizing disoccluded or out-of-view content. We identify Evid…

DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-based Human Video Generation

2024-01-01 · CVPR 2024 1 · Chenyang Wang, Zerong Zheng, Tao Yu, Xiaoqian Lv 외

Existing diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In …

Video Generation