paper-with-me

홈 › Papers

RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation

2026-08-28 · Zhen Xiao, Zhen Shen, Zhaofan Qiu, Ting Yao, Xueliang Liu, Tao Mei arxiv

Cross-identity character animation aims to drive a target identity from a reference image to follow the motion of a source character from a driving video. The core challenge lies in the inherent entanglement of two capabilities: cross-identity spatial mapping (aligning position, scale, and skeletal proportions) and motion control (refining joint articulation, volumetric consistency, and view coherence). We introduce Reference-Aware Structural Alignment (RASA), a framework that disentangles spatial mapping from motion control by injecting structured priors into a Diffusion Transformer (DiT). Our approach has two stages. First, a Spatial Prior Calibrator (SPC) fuses reference identity with driving pose to generate a spatially grounded initial noise latent, ensuring correct positioning, scaling, and alignment with the driving skeleton. Second, an Inherent Motional Guider (IMG) encodes shape-agnostic SMPL articulation parameters into a semantic motion vector beyond appearance-biased 2D keypoints. Injected into intermediate DiT layers, this vector complements the base pose condition for anatomically consistent articulation and view-aware volumetric refinement. We curate CIM-Bench, a high-quality benchmark with rigorous curation, for evaluation. Extensive experiments show RASA significantly outperforms state-of-the-art methods in motion fidelity and visual quality. Our work establishes a new paradigm showing disentangled spatial and motional priors are key to robust character animation. Project page: https://hidream.ai.github.io/RASA/

📄 PDF Abstract BibTeX arXiv:2608.28219

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Characterizing functional brain networks and emotional centers based on Rasa theory of Indian aesthetics

2018-09-14

In Indian history of arts, Rasas are the aesthetics associated with any auditory, visual, literary or musical piece of art that evokes highly orchestrated emotional states. In this work, we study the functional response …

Community DetectionEEGElectroencephalogram (EEG)Functional Connectivity

Tropical Land Use Land Cover Mapping in Pará (Brazil) using Discriminative Markov Random Fields and Multi-temporal TerraSAR-X Data

2017-09-22 · Ron Hagensieker, Ribana Roscher, Johannes Rosentreter, Benjamin Jakimow 외

Remote sensing satellite data offer the unique possibility to map land use land cover transformations by providing spatially explicit information. However, detection of short-term processes and land use patterns of high …

ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism

2024-09-05 · Hsing-Hang Chou, Yun-Shao Lin, Ching-Chin Sung, Yu Tsao 외

The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving appl…

Emotion ClassificationVoice Conversion

Disentangled Textual Priors for Diffusion-based Image Super-Resolution

2026-03-08 · Lei Jiang, Xin Liu, Xinze Tong, Zhiliang Li 외 arxiv

Image Super-Resolution (SR) aims to reconstruct high-resolution images from degraded low-resolution inputs. While diffusion-based SR methods offer powerful generative capabilities, their performance heavily depends on ho…

Image Super-Resolution

Orthogonal Spatial-temporal Distributional Transfer for 4D Generation

2026-03-05 · Wei Liu, Shengqiong Wu, Bobo Li, Haoyu Zhao 외 arxiv

In the AIGC era, generating high-quality 4D content has garnered increasing research attention. Unfortunately, current 4D synthesis research is severely constrained by the lack of large-scale 4D datasets, preventing mode…