paper-with-me

홈 › Papers

RealisDance-DiT: Simple yet Strong Baseline towards Controllable Character Animation in the Wild

2025-04-21 · Jingkai Zhou, Yifan Wu, Shikai Li, Min Wei, Chao Fan, Weihua Chen, Wei Jiang, Fan Wang

Controllable character animation remains a challenging problem, particularly in handling rare poses, stylized characters, character-object interactions, complex illumination, and dynamic scenes. To tackle these issues, prior work has largely focused on injecting pose and appearance guidance via elaborate bypass networks, but often struggles to generalize to open-world scenarios. In this paper, we propose a new perspective that, as long as the foundation model is powerful enough, straightforward model modifications with flexible fine-tuning strategies can largely address the above challenges, taking a step towards controllable character animation in the wild. Specifically, we introduce RealisDance-DiT, built upon the Wan-2.1 video foundation model. Our sufficient analysis reveals that the widely adopted Reference Net design is suboptimal for large-scale DiT models. Instead, we demonstrate that minimal modifications to the foundation model architecture yield a surprisingly strong baseline. We further propose the low-noise warmup and "large batches and small iterations" strategies to accelerate model convergence during fine-tuning while maximally preserving the priors of the foundation model. In addition, we introduce a new test dataset that captures diverse real-world challenges, complementing existing benchmarks such as TikTok dataset and UBC fashion video dataset, to comprehensively evaluate the proposed method. Extensive experiments show that RealisDance-DiT outperforms existing methods by a large margin.

📄 PDF Abstract BibTeX arXiv:2504.14977

Code (1)

damo-cv/realisdance pytorch

Similar Papers 제목 키워드 기반

RealisDance: Equip controllable character animation with realistic hands

2024-09-10 · Jingkai Zhou, Benzhi Wang, Weihua Chen, Jingqi Bai 외

Controllable character animation is an emerging task that generates character videos controlled by pose sequences from given character images. Although character consistency has made significant progress via reference UN…

On Variational Learning of Controllable Representations for Text without Supervision

2019-05-28 · ICML 2020 1 · Peng Xu, Jackie Chi Kit Cheung, Yanshuai Cao

The variational autoencoder (VAE) can learn the manifold of natural images on certain datasets, as evidenced by meaningful interpolating or extrapolating in the continuous latent space. However, on discrete data such as …

Style TransferText GenerationText Style Transfer

MUGL: Large Scale Multi Person Conditional Action Generation with Locomotion

2021-10-21 · Shubh Maheshwari, Debtanu Gupta, Ravi Kiran Sarvadevabhatla

We introduce MUGL, a novel deep neural model for large-scale, diverse generation of single and multi-person pose-based action sequences with locomotion. Our controllable approach enables variable-length generations custo…

Action GenerationDiversityHuman action generation

Controllable Molecular Generative Foundation Models

2026-05-14 · Yihan Zhu, Yuhan Liu, Weijiang Li, Tengfei Luo 외 arxiv

Despite the success of foundation models in language and vision, molecular graph generation still lacks a unified framework for heterogeneous design tasks with reliable controllability. While reinforcement learning (RL) …

Reinforcement LearningGraph GenerationDrug Discovery

STRONG -- Structure Controllable Legal Opinion Summary Generation

2023-09-29 · Yang Zhong, Diane Litman

We propose an approach for the structure controllable summarization of long legal opinions that considers the argument structure of the document. Our approach involves using predicted argument role information to guide t…