paper-with-me

Papers

Optimizing ID Consistency in Multimodal Large Models: Facial Restoration via Alignment, Entanglement, and Disentanglement

2026-02-21 · Yuran Dong, Hang Dai, Mang Ye arxiv

Multimodal editing large models have demonstrated powerful editing capabilities across diverse tasks. However, a persistent and long-standing limitation is the decline in facial identity (ID) consistency during realistic portrait editing. Due to the human eye's high sensitivity to facial features, such inconsistency significantly hinders the practical deployment of these models. Current facial ID preservation methods struggle to achieve consistent restoration of both facial identity and edited element IP due to Cross-source Distribution Bias and Cross-source Feature Contamination. To address these issues, we propose EditedID, an Alignment-Disentanglement-Entanglement framework for robust identity-specific facial restoration. By systematically analyzing diffusion trajectories, sampler behaviors, and attention properties, we introduce three key components: 1) Adaptive mixing strategy that aligns cross-source latent representations throughout the diffusion process. 2) Hybrid solver that disentangles source-specific identity attributes and details. 3) Attentional gating mechanism that selectively entangles visual elements. Extensive experiments show that EditedID achieves state-of-the-art performance in preserving original facial ID and edited element IP consistency. As a training-free and plug-and-play solution, it establishes a new benchmark for practical and reliable single/multi-person facial identity restoration in open-world settings, paving the way for the deployment of multimodal editing large models in real-person editing scenarios. The code is available at https://github.com/NDYBSNDY/EditedID.

📄 PDF Abstract BibTeX arXiv:2602.18752

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration

2026-04-16 · Zheng Chen, Bowen Chai, Rongjun Gao, Mingtao Nie 외 arxiv

Video face restoration aims to enhance degraded face videos into high-quality results with realistic facial details, stable identity, and temporal coherence. Recent diffusion-based methods have brought strong generative …

RGFVR: Reference-Guided Face Video Restoration with Flow Matching

2026-06-15 · Cem Eteke, Batuhan Tosun, Eckehard Steinbach arxiv

Face video restoration from degraded observations is challenging, as it requires simultaneously recovering visual fidelity, temporal consistency, and subject identity. Existing approaches are often either reference-free,…

Video Restoration

RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models

2024-07-25 · Haoyu Chen, Wenbo Li, Jinjin Gu, Jingjing Ren 외

Natural images captured by mobile devices often suffer from multiple types of degradation, such as noise, blur, and low light. Traditional image restoration methods require manual selection of specific tasks, algorithms,…

Image RestorationLow-Light Image Enhancement

ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving

2024-04-25 · Jiehui Huang, Xiao Dong, Wenhui Song, Zheng Chong 외

Diffusion-based technologies have made significant strides, particularly in personalized and customized facialgeneration. However, existing methods face challenges in achieving high-fidelity and detailed identity (ID)con…

Diversity

TANet: A new Paradigm for Global Face Super-resolution via Transformer-CNN Aggregation Network

2021-09-16 · Yuanzhi Wang, Tao Lu, Yanduo Zhang, Junjun Jiang 외

Recently, face super-resolution (FSR) methods either feed whole face image into convolutional neural networks (CNNs) or utilize extra facial priors (e.g., facial parsing maps, facial landmarks) to focus on facial structu…

Face ReconstructionSuper-Resolution