paper-with-me

홈 › Papers

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

2025-07-02 · Shuai Tan, Bill Gong, Bin Ji, Ye Pan arxiv

Talking head generation is gaining significant importance across various domains, with a growing demand for high-quality rendering. However, existing methods often suffer from identity leakage (IL) and rendering artifacts (RA), particularly in extreme cases. Through an in-depth analysis of previous approaches, we identify two key insights: (1) IL arises from identity information embedded within motion features, and (2) this identity information can be leveraged to address RA. Building on these findings, this paper introduces FixTalk, a novel framework designed to simultaneously resolve both issues for high-quality talking head generation. Firstly, we propose an Enhanced Motion Indicator (EMI) to effectively decouple identity information from motion features, mitigating the impact of IL on generated talking heads. To address RA, we introduce an Enhanced Detail Indicator (EDI), which utilizes the leaked identity information to supplement missing details, thus fixing the artifacts. Extensive experiments demonstrate that FixTalk effectively mitigates IL and RA, achieving superior performance compared to state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2507.01390

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head Generation

Similar Papers 제목 키워드 기반

IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation

2025-12-29 · Donghao Zhou, Jingyu Lin, Guibao Shen, Quande Liu 외 arxiv

Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining detailed and diverse human face consistenc…

Story Generation

Taming Identity Consistency and Prompt Diversity in Diffusion Models via Latent Concatenation and Masked Conditional Flow Matching

2025-11-11 · Aditi Singhania, Arushi Jain, Krutik Malani, Riddhi Dhawan 외 arxiv

Subject-driven image generation aims to synthesize novel depictions of a specific subject across diverse contexts while preserving its core identity features. Achieving both strong identity consistency and high prompt di…

parameter-efficient fine-tuningImage Generation

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration

2026-06-23 · Yang Zhou, Wenxue Li, Peng Zhang, Yifei Chen 외 arxiv

Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. Existing methods often struggle to simultaneously address thr…

Video RestorationVideo Generation

CopyCat: Many-to-Many Fine-Grained Prosody Transfer for Neural Text-to-Speech

2020-04-30

Prosody Transfer (PT) is a technique that aims to use the prosody from a source audio as a reference while synthesising speech. Fine-grained PT aims at capturing prosodic aspects like rhythm, emphasis, melody, duration, …

Rhythmtext-to-speechText to Speech

Evaluating Identity Leakage in Speaker De-Identification Systems

2025-08-19 · Seungmin Seo, Oleg Aulov, Afzal Godil, Kevin Mangold arxiv

Speaker de-identification aims to conceal a speaker's identity while preserving intelligibility of the underlying speech. We introduce a benchmark that quantifies residual identity leakage with three complementary error …