paper-with-me

Papers

Overcoming False Illusions in Real-World Face Restoration with Multi-Modal Guided Diffusion Model

2024-10-05 · Keda Tao, Jinjin Gu, Yulun Zhang, Xiucheng Wang, Nan Cheng

We introduce a novel Multi-modal Guided Real-World Face Restoration (MGFR) technique designed to improve the quality of facial image restoration from low-quality inputs. Leveraging a blend of attribute text prompts, high-quality reference images, and identity information, MGFR can mitigate the generation of false facial attributes and identities often associated with generative face restoration methods. By incorporating a dual-control adapter and a two-stage training strategy, our method effectively utilizes multi-modal prior information for targeted restoration tasks. We also present the Reface-HQ dataset, comprising over 21,000 high-resolution facial images across 4800 identities, to address the need for reference face training images. Our approach achieves superior visual quality in restoring facial details under severe degradation and allows for controlled restoration processes, enhancing the accuracy of identity preservation and attribute correction. Including negative quality samples and attribute prompts in the training further refines the model's ability to generate detailed and perceptually accurate images.

📄 PDF Abstract BibTeX arXiv:2410.04161

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage Restoration

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

Towards Illusions Awareness in Cyber-Physical System's Design

2026-09-15 · Anna Di Placido, Nicolas Ferry, Julien Deantoni arxiv

Cyber-Physical Systems (CPS) operate through a continuous sense-compute-act loop within an open context environment, making it impossible to anticipate all the situations the system will face. To cope with this openness,…

IllusionBench: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models

2025-01-01 · Yiming Zhang, ZiCheng Zhang, Xinyi Wei, Xiaohong Liu 외

Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have bee…

HallucinationMultiple-choice

Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?

2023-10-31 · Yichi Zhang, Jiayi Pan, Yuchen Zhou, Rui Pan 외

Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human's perception of reality isn't always faithful to th…

LanEvil: Benchmarking the Robustness of Lane Detection to Environmental Illusions

2024-06-03 · Tianyuan Zhang, Lu Wang, Hainan Li, Yisong Xiao 외

Lane detection (LD) is an essential component of autonomous driving systems, providing fundamental functionalities like adaptive cruise control and automated lane centering. Existing LD benchmarks primarily focus on eval…

Autonomous DrivingBenchmarkingLane Detection

Diffusion Illusions: Hiding Images in Plain Sight

2023-12-06 · Ryan Burgert, Xiang Li, Abe Leite, Kanchana Ranasinghe 외

We explore the problem of computationally generating special `prime' images that produce optical illusions when physically arranged and viewed in a certain way. First, we propose a formal definition for this problem. Nex…