paper-with-me

홈 › Papers

High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer

2025-12-26 · Shen Zheng, Jiaran Cai, Yuansheng Guan, Shenneng Huang, Xingpei Ma, Junjie Cao, Hanfeng Zhao, Qiang Zhang, Shunsi Zhang, Xiao-Ping Zhang arxiv

Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or regular motions, significant challenges remain, particularly in generating long-duration videos. Furthermore, the synthesis of fine-grained facial and hand details remains under-explored, limiting the applicability of current approaches in real-world, high-quality applications. To address these limitations, we propose a diffusion transformer (DiT)-based framework which focuses on generating high-fidelity and long-duration human animation videos. First, we design a set of hybrid implicit guidance signals and a sharpness guidance factor, enabling our framework to additionally incorporate detailed facial and hand features as guidance. Next, we incorporate the time-aware position shift fusion module, modify the input format within the DiT backbone, and refer to this mechanism as the Position Shift Adaptive Module, which enables video generation of arbitrary length. Finally, we introduce a novel data augmentation strategy and a skeleton alignment model to reduce the impact of human shape variations across different identities. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches, achieving superior performance in both high-fidelity and long-duration human image animation.

📄 PDF Abstract BibTeX arXiv:2512.21905

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationVideo Generation

Similar Papers 제목 키워드 기반

PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation

2025-08-07 · Jingxuan He, Busheng Su, Finn Wong arxiv

Generating temporally coherent, long-duration videos with precise control over subject identity and movement remains a fundamental challenge for contemporary diffusion-based models, which often suffer from identity drift…

Video Generation

DiVa-360: The Dynamic Visual Dataset for Immersive Neural Fields

2023-07-31 · CVPR 2024 1 · Cheng-You Lu, Peisen Zhou, Angela Xing, Chandradeep Pokhariya 외

Advances in neural fields are enabling high-fidelity capture of the shape and appearance of dynamic 3D scenes. However, their capabilities lag behind those offered by conventional representations such as 2D videos becaus…

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

2026-04-11 · Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan arxiv

Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantica…

Video Summarization

Hierarchical Semi-Markov Models with Duration-Aware Dynamics for Activity Sequences

2025-09-22 · Rohit Dube, Natarajan Gautam, Amarnath Banerjee, Harsha Nagarajan arxiv

Residential electricity demand at granular scales is driven by what people do and for how long. Accurately forecasting this demand for applications like microgrid management and demand response therefore requires generat…

Magnetically Self-Sealed MR Haptic Actuator With PWM-Based Excitation and High-Fidelity Torque Control

2026-08-20 · Dong Qiang, Tian Yuan, Song Yang, Kequan Xia 외 arxiv

Accurate and stable torque rendering is essential for safe and perceptive human--machine interaction. Magnetorheological fluid (MRF)-based actuators offer a compact and rapidly controllable solution for haptic feedback, …