paper-with-me

홈 › Papers

PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence

2025-12-15 · Ruiyan Wang, Teng Hu, Kaihui Huang, Zihan Su, Ran Yi, Lizhuang Ma arxiv

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current pose-guided video generation methods are limited to accepting only human poses as input, thus generalizing poorly to pose of other subjects. To address this issue, we propose PoseAnything, the first universal pose-guided video generation framework capable of handling both human and non-human characters, supporting arbitrary skeletal inputs. To enhance consistency preservation during motion, we introduce Part-aware Temporal Coherence Module, which divides the subject into different parts, establishes part correspondences, and computes cross-attention between corresponding parts across frames to achieve fine-grained part-level consistency. Additionally, we propose Subject and Camera Motion Decoupled CFG, a novel guidance strategy that, for the first time, enables independent camera movement control in pose-guided video generation, by separately injecting subject and camera motion control information into the positive and negative anchors of CFG. Furthermore, we present XPose, a high-quality public dataset containing 50,000 non-human pose-video pairs, along with an automated pipeline for annotation and filtering. Extensive experiments demonstrate that Pose-Anything significantly outperforms state-of-the-art methods in both effectiveness and generalization.

📄 PDF Abstract BibTeX arXiv:2512.13465

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

ComposeAnything: Composite Object Priors for Text-to-Image Generation

2025-05-30 · Zeeshan Khan, ShiZhe Chen, Cordelia Schmid

Generating images from text involving complex and novel object arrangements remains a significant challenge for current text-to-image (T2I) models. Although prior layout-based methods improve object arrangements using sp…

DenoisingImage GenerationObjectText to Image Generation+1

ARDuP: Active Region Video Diffusion for Universal Policies

2024-06-19 · Shuaiyi Huang, Mara Levy, Zhenyu Jiang, Anima Anandkumar 외

Sequential decision-making can be formulated as a text-conditioned video generation problem, where a video planner, guided by a text-defined goal, generates future frames visualizing planned actions, from which control a…

Decision MakingSequential Decision MakingVideo Generation

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation

2025-02-01 · Yang Cao, Zhao Song, Chiwun Yang

This paper considers an efficient video modeling process called Video Latent Flow Matching (VLFM). Unlike prior works, which randomly sampled latent patches for video generation, our method relies on current strong pre-t…

Image GenerationVideo Generation

Learning Universal Policies via Text-Guided Video Generation

2023-01-31 · NeurIPS 2023 11

A goal of artificial intelligence is to construct an agent that can solve a wide variety of tasks. Recent progress in text-guided image synthesis has yielded models with an impressive ability to generate complex novel im…

Decision MakingImage GenerationRobot ManipulationSequential Decision Making+2

EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation

2024-08-23 · Cong Wang, Jiaxi Gu, Panwen Hu, Haoyu Zhao 외

Following the advancements in text-guided image generation technology exemplified by Stable Diffusion, video generation is gaining increased attention in the academic community. However, relying solely on text guidance f…

Image GenerationVideo Generation