paper-with-me

홈 › Papers

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

2026-03-15 · Ngoc-Son Nguyen, Thanh V. T. Tran, Jeongsoo Choi, Hieu-Nghia Huynh-Nguyen, Truong-Son Hy, Van Nguyen arxiv

Video dubbing requires content accuracy, expressive prosody, high-quality acoustics, and precise lip synchronization, yet existing approaches struggle on all four fronts. To address these issues, we propose DiFlowDubber, the first video dubbing framework built upon a discrete flow matching backbone with a novel two-stage training strategy. In the first stage, a zero-shot text-to-speech (TTS) system is pre-trained on large-scale corpora, where a deterministic architecture captures linguistic structures, and the Discrete Flow-based Prosody-Acoustic (DFPA) module models expressive prosody and realistic acoustic characteristics. In the second stage, we propose the Content-Consistent Temporal Adaptation (CCTA) to transfer TTS knowledge to the dubbing domain: its Synchronizer enforces cross-modal alignment for lip-synchronized speech. Complementarily, the Face-to-Prosody Mapper (FaPro) conditions prosody on facial expressions, whose outputs are then fused with those of the Synchronizer to construct rich, fine-grained multimodal embeddings that capture prosody-content correlations, guiding the DFPA to generate expressive prosody and acoustic tokens for content-consistent speech. Experiments on two benchmark datasets demonstrate that DiFlowDubber outperforms prior methods across multiple evaluation metrics.

📄 PDF Abstract BibTeX arXiv:2603.14267

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Discrete Flow Matching

2024-07-22 · Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk 외

Despite Flow Matching and diffusion models having emerged as powerful generative paradigms for continuous variables such as images and videos, their application to high-dimensional discrete data, such as language, is sti…

HumanEvalmbppPrediction

Fisher Flow Matching for Generative Modeling over Discrete Data

2024-05-23 · Oscar Davis, Samuel Kessler, Mircea Petrache, İsmail İlkan Ceylan 외

Generative modeling over discrete data has recently seen numerous success stories, with applications spanning language modeling, biological sequence design, and graph-structured molecular data. The predominant generative…

Language ModelingLanguage ModellingVideo Generation

MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation

2025-02-16 · Michael Fuest, Vincent Tao Hu, Björn Ommer

Generating long, high-quality videos remains a challenge due to the complex interplay of spatial and temporal dynamics and hardware limitations. In this work, we introduce \textbf{MaskFlow}, a unified video generation fr…

Video Generation

Tessellations of Semi-Discrete Flow Matching

2026-05-08 · Emile Pierret, Johannes Hertrich, Samuel Hurault, Julie Delon arxiv

We study Flow Matching in a semi-discrete setting where a Gaussian source is transported toward a discrete target supported on finitely many points. This semi-discrete regime is the theoretical setting behind the use of …

Flowception: Temporally Expansive Flow Matching for Video Generation

2025-12-12 · Tariq Berrada Ifriqi, John Nguyen, Karteek Alahari, Jakob Verbeek 외 arxiv

We present Flowception, a novel non-autoregressive and variable-length video generation framework. Flowception learns a probability path that interleaves discrete frame insertions with continuous frame denoising. Compare…

Video Generation