paper-with-me

홈 › Papers

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

2026-05-28 · Hao Wu, Xiangyang Luo, Hao Wang, Jiawei Zhang, Yi Zhang, Jinwei Wang arxiv

With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets, resulting in high computational costs that hinder scalability and accessibility of diffusion-based approaches across the research community. To address this, we propose a finetuning-free paradigm that directly performs talking face generation using the pretrained weights of Stable Diffusion and IP-Adapter. This backbone leverages the visual embedding capability of IP-Adapter to mine lip-related semantics from the pretrained Stable Diffusion. To address the challenges of identity drift, synchronization errors, and temporal instability, we also design three trainable-parameterfree components: (1) the Structurist, which explicitly disentangles and reassembles lip and appearance features to mitigate identity drift and appearance distortion; (2) the Structure Controller, which adaptively refines embeddings based on quasi-monotonic motion trends for precise lip synchronization; and (3) the Noise Sensor, which introduces Gaussian prior to detect and suppress flicker and jitter artifacts and enhance temporal consistency. Experimental results show that our method outperforms existing SOTA approaches in both lip-sync accuracy (at least 0.16 gain in PCLD) and visual fidelity (at least 0.7 improvement in FID), establishing a novel fine-tuning-free diffusion framework for talking face generation.

📄 PDF Abstract BibTeX arXiv:2605.30230

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Face Generation

Similar Papers 제목 키워드 기반

LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation

2025-01-27 · Farzad Farhadzadeh, Debasmit Das, Shubhankar Borse, Fatih Porikli

The rising popularity of large foundation models has led to a heightened demand for parameter-efficient fine-tuning methods, such as Low-Rank Adaptation (LoRA), which offer performance comparable to full model fine-tunin…

Image Generationparameter-efficient fine-tuningText to Image GenerationText-to-Image Generation

MaTe: Images Are All You Need for Material Transfer via Diffusion Transformer

2026-05-15 · Nisha Huang, Henglin Liu, Yizhou Lin, Kaer Huang 외 arxiv

Recent diffusion-based methods for material transfer rely on image fine-tuning or complex architectures with assistive networks, but face challenges including text dependency, extra computational costs, and feature misal…

Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance

2025-09-26 · Luc Boudier, Loris Manganelli, Eleftherios Tsonis, Nicolas Dufour 외 arxiv

Few-shot image classification remains challenging due to the limited availability of labeled examples. Recent approaches have explored generating synthetic training data using text-to-image diffusion models, but often re…

Few-Shot Image ClassificationImage-to-Image TranslationSynthetic Data Generation

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

2023-08-13 · Hu Ye, Jun Zhang, Sibo Liu, Xiao Han 외

Recent years have witnessed the strong power of large text-to-image diffusion models for the impressive generative capability to create high-fidelity images. However, it is very tricky to generate desired images using on…

Diffusion Personalization Tuning FreeImage GenerationPersonalized Image GenerationPrompt Engineering

OrthoFuse: Training-free Riemannian Fusion of Orthogonal Style-Concept Adapters for Diffusion Models

2026-04-06 · Ali Aliev, Kamil Garifullin, Nikolay Yudin, Vera Soboleva 외 arxiv

In a rapidly growing field of model training there is a constant practical interest in parameter-efficient fine-tuning and various techniques that use a small amount of training data to adapt the model to a narrow task. …

parameter-efficient fine-tuning