paper-with-me

홈 › Papers

PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control

2024-05-23 · Yong Zhong, Min Zhao, Zebin You, Xiaofeng Yu, Changwang Zhang, Chongxuan Li

In this paper, we introduce PoseCrafter, a one-shot method for personalized video generation following the control of flexible poses. Built upon Stable Diffusion and ControlNet, we carefully design an inference process to produce high-quality videos without the corresponding ground-truth frames. First, we select an appropriate reference frame from the training video and invert it to initialize all latent variables for generation. Then, we insert the corresponding training pose into the target pose sequences to enhance faithfulness through a trained temporal attention module. Furthermore, to alleviate the face and hand degradation resulting from discrepancies between poses of training videos and inference poses, we implement simple latent editing through an affine transformation matrix involving facial and hand landmarks. Extensive experiments on several datasets demonstrate that PoseCrafter achieves superior results to baselines pre-trained on a vast collection of videos under 8 commonly used metrics. Besides, PoseCrafter can follow poses from different individuals or artificial edits and simultaneously retain the human identity in an open-domain training video. Our project page is available at https://ml-gsai.github.io/PoseCrafter-demo/.

📄 PDF Abstract BibTeX arXiv:2405.14582

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis

2025-10-22 · Qing Mao, Tianxin Huang, Yu Zhu, Jinqiu Sun 외 arxiv

Pairwise camera pose estimation from sparsely overlapping image pairs remains a critical and unsolved challenge in 3D vision. Most existing methods struggle with image pairs that have small or no overlap. Recent approach…

Camera Pose EstimationNovel View SynthesisVideo Generation

Zero-shot personalized lip-to-speech synthesis with face image based voice control

2023-05-09 · Zheng-Yan Sheng, Yang Ai, Zhen-Hua Ling

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. Howev…

Lip to Speech SynthesisRepresentation LearningSpeech Synthesis

Lynx: Towards High-Fidelity Personalized Video Generation

2025-09-19 · Shen Sang, Tiancheng Zhi, Tianpei Gu, Jing Liu 외 arxiv

We present Lynx, a high-fidelity model for personalized video synthesis from a single input image. Built on an open-source Diffusion Transformer (DiT) foundation model, Lynx introduces two lightweight adapters to ensure …

Video Generation

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

2026-04-19 · Junjia Huang, Binbin Yang, Pengxiang Yan, Jiyang Liu 외 arxiv

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing a…

Visual StorytellingStory Continuation

CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers

2025-02-10 · D. She, Mushui Liu, Jingxuan Pang, Jin Wang 외

Customized generation has achieved significant progress in image synthesis, yet personalized video generation remains challenging due to temporal inconsistencies and quality degradation. In this paper, we introduce Custo…

Image GenerationVideo Generation