paper-with-me

홈 › Papers

MoCha:End-to-End Video Character Replacement without Structural Guidance

2026-01-13 · Zhengbo Xu, Jie Ma, Ziheng Wang, Zhan Peng, Jun Liang, Jing Li arxiv

Controllable video character replacement with a user-provided identity remains a challenging problem due to the lack of paired video data. Prior works have predominantly relied on a reconstruction-based paradigm that requires per-frame segmentation masks and explicit structural guidance (e.g., skeleton, depth). This reliance, however, severely limits their generalizability in complex scenarios involving occlusions, character-object interactions, unusual poses, or challenging illumination, often leading to visual artifacts and temporal inconsistencies. In this paper, we propose MoCha, a pioneering framework that bypasses these limitations by requiring only a single arbitrary frame mask. To effectively adapt the multi-modal input condition and enhance facial identity, we introduce a condition-aware RoPE and employ an RL-based post-training stage. Furthermore, to overcome the scarcity of qualified paired-training data, we propose a comprehensive data construction pipeline. Specifically, we design three specialized datasets: a high-fidelity rendered dataset built with Unreal Engine 5 (UE5), an expression-driven dataset synthesized by current portrait animation techniques, and an augmented dataset derived from existing video-mask pairs. Extensive experiments demonstrate that our method substantially outperforms existing state-of-the-art approaches. We will release the code to facilitate further research. Please refer to our project page for more details: orange-3dv-team.github.io/MoCha

📄 PDF Abstract BibTeX arXiv:2601.08587

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoCha: Towards Movie-Grade Talking Character Synthesis

2025-03-30 · Cong Wei, Bo Sun, Haoyu Ma, Ji Hou 외

Recent advancements in video generation have achieved impressive motion realism, yet they often overlook character-driven storytelling, a crucial task for automated film, animation generation. We introduce Talking Charac…

Video Generation

MOCHA: Discovering Multi-Order Dynamic Causality in Temporal Point Processes

2025-08-26 · Yunyang Cao, Juekai Lin, Wenhao Li, Bo Jin arxiv

Discovering complex causal dependencies in temporal point processes (TPPs) is critical for modeling real-world event sequences. Existing methods typically rely on static or first-order causal structures, overlooking the …

Point Processes

MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing

2025-08-20 · Jeahun Sung, Changhyun Roh, Chanho Eom, Jihyong Oh arxiv

Recent advances in portable imaging have made camera-based screen capture ubiquitous. Unfortunately, frequency aliasing between the camera's color filter array (CFA) and the display's sub-pixels induces moiré patterns th…

Multi-head Monotonic Chunkwise Attention For Online Speech Recognition

2020-05-01 · Baiji Liu, Songjun Cao, Sining Sun, Weibin Zhang 외

The attention mechanism of the Listen, Attend and Spell (LAS) model requires the whole input sequence to calculate the attention context and thus is not suitable for online speech recognition. To deal with this problem, …

speech-recognitionSpeech Recognition

A comparison of streaming models and data augmentation methods for robust speech recognition

2021-11-19 · Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg 외

In this paper, we present a comparative study on the robustness of two different online streaming speech recognition models: Monotonic Chunkwise Attention (MoChA) and Recurrent Neural Network-Transducer (RNN-T). We explo…

Data AugmentationRobust Speech Recognitionspeech-recognitionSpeech Recognition