paper-with-me

홈 › Papers

Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation

2024-05-26 · Jinlin Liu, Kai Yu, Mengyang Feng, Xiefan Guo, Miaomiao Cui

Recent advancements in human video synthesis have enabled the generation of high-quality videos through the application of stable diffusion models. However, existing methods predominantly concentrate on animating solely the human element (the foreground) guided by pose information, while leaving the background entirely static. Contrary to this, in authentic, high-quality videos, backgrounds often dynamically adjust in harmony with foreground movements, eschewing stagnancy. We introduce a technique that concurrently learns both foreground and background dynamics by segregating their movements using distinct motion representations. Human figures are animated leveraging pose-based motion, capturing intricate actions. Conversely, for backgrounds, we employ sparse tracking points to model motion, thereby reflecting the natural interaction between foreground activity and environmental changes. Training on real-world videos enhanced with this innovative motion depiction approach, our model generates videos exhibiting coherent movement in both foreground subjects and their surrounding contexts. To further extend video generation to longer sequences without accumulating errors, we adopt a clip-by-clip generation strategy, introducing global features at each step. To ensure seamless continuity across these segments, we ingeniously link the final frame of a produced clip with input noise to spawn the succeeding one, maintaining narrative flow. Throughout the sequential generation process, we infuse the feature representation of the initial reference image into the network, effectively curtailing any cumulative color inconsistencies that may otherwise arise. Empirical evaluations attest to the superiority of our method in producing videos that exhibit harmonious interplay between foreground actions and responsive background dynamics, surpassing prior methodologies in this regard.

📄 PDF Abstract BibTeX arXiv:2405.16393

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation

2025-10-01 · Yunbo Xu, Xuesong Zhang, Jia Li, Zhenzhen Hu 외 arxiv

Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has propelled advancements in VLN, the signifi…

Vision-Language Navigation

Disentangling Motion, Foreground and Background Features in Videos

2017-07-13 · Xunyu Lin, Victor Campos, Xavier Giro-i-Nieto, Jordi Torres 외

This paper introduces an unsupervised framework to extract semantically rich features for video representation. Inspired by how the human visual system groups objects based on motion cues, we propose a deep convolutional…

OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance

2026-03-19 · Cong Wang, Hanxin Zhu, Xiao Tang, Jiayi Luo 외 arxiv

Recent progress in video generation has led to substantial improvements in visual fidelity, yet ensuring physically consistent motion remains a fundamental challenge. Intuitively, this limitation can be attributed to the…

Video Generation

Dance Dance Generation: Motion Transfer for Internet Videos

2019-03-30 · Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui 외

This work presents computational methods for transferring body movements from one person to another with videos collected in the wild. Specifically, we train a personalized model on a single video from the Internet which…

Intrinsic Harmonization for Illumination-Aware Compositing

2023-12-06 · Chris Careaga, S. Mahdi H. Miangoleh, Yağız Aksoy

Despite significant advancements in network-based image harmonization techniques, there still exists a domain disparity between typical training pairs and real-world composites encountered during inference. Most existing…

Image HarmonizationImage Relighting