paper-with-me

Papers

DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer

2026-01-04 · Xu Guo, Fulong Ye, Xinghui Li, Pengqi Tu, Pengze Zhang, Qichao Sun, Songtao Zhao, Xiangwang Hou, Qian He arxiv

Video Face Swapping (VFS) requires seamlessly injecting a source identity into a target video while meticulously preserving the original pose, expression, lighting, background, and dynamic information. Existing methods struggle to maintain identity similarity and attribute preservation while preserving temporal consistency. To address the challenge, we propose a comprehensive framework to seamlessly transfer the superiority of Image Face Swapping (IFS) to the video domain. We first introduce a novel data pipeline SyncID-Pipe that pre-trains an Identity-Anchored Video Synthesizer and combines it with IFS models to construct bidirectional ID quadruplets for explicit supervision. Building upon paired data, we propose the first Diffusion Transformer-based framework DreamID-V, employing a core Modality-Aware Conditioning module to discriminatively inject multi-model conditions. Meanwhile, we propose a Synthetic-to-Real Curriculum mechanism and an Identity-Coherence Reinforcement Learning strategy to enhance visual realism and identity consistency under challenging scenarios. To address the issue of limited benchmarks, we introduce IDBench-V, a comprehensive benchmark encompassing diverse scenes. Extensive experiments demonstrate DreamID-V outperforms state-of-the-art methods and further exhibits exceptional versatility, which can be seamlessly adapted to various swap-related tasks.

📄 PDF Abstract BibTeX arXiv:2601.01425

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningFace Swapping

Similar Papers 제목 키워드 기반

DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning

2025-04-20 · Fulong Ye, Miao Hua, Pengze Zhang, Xinghui Li 외

In this paper, we introduce DreamID, a diffusion-based face swapping model that achieves high levels of ID similarity, attribute preservation, image fidelity, and fast inference speed. Unlike the typical face swapping tr…

AttributeFace SwappingTriplet

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

2026-02-12 · Xu Guo, Fulong Ye, Qichao Sun, Liyang Chen 외 arxiv

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video e…

Video Generation

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models

2026-05-29 · Jiazheng Xing, Hangjie Yuan, Lingling Cai, Xinyu Liu 외 arxiv

Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large high-fidelity generator into the unified training loop is computationally prohibit…

Video Generation

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

2026-06-10 · Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin 외 arxiv

Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but directly using them for robot learning …

Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging

2025-03-28 · Chongjie Ye, Yushuang Wu, Ziteng Lu, Jiahao Chang 외

With the growing demand for high-fidelity 3D models from 2D images, existing methods still face significant challenges in accurately reproducing fine-grained geometric details due to limitations in domain gaps and inhere…

3D geometry