paper-with-me

홈 › Papers

ReFIR: Grounding Large Restoration Models with Retrieval Augmentation

2024-10-08 · Hang Guo, Tao Dai, Zhihao Ouyang, Taolin Zhang, Yaohua Zha, Bin Chen, Shu-Tao Xia

Recent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents or textures when dealing with severe degradations, due to their heavy reliance on limited internal knowledge. In this paper, we propose an orthogonal solution called the Retrieval-augmented Framework for Image Restoration (ReFIR), which incorporates retrieved images as external knowledge to extend the knowledge boundary of existing LRMs in generating details faithful to the original scene. Specifically, we first introduce the nearest neighbor lookup to retrieve content-relevant high-quality images as reference, after which we propose the cross-image injection to modify existing LRMs to utilize high-quality textures from retrieved images. Thanks to the additional external knowledge, our ReFIR can well handle the hallucination challenge and facilitate faithfully results. Extensive experiments demonstrate that ReFIR can achieve not only high-fidelity but also realistic restoration results. Importantly, our ReFIR requires no training and is adaptable to various LRMs.

📄 PDF Abstract BibTeX arXiv:2410.05601

Code (1)

csguoh/refir 공식 구현 pytorch

Tasks

HallucinationImage RestorationRetrieval

Similar Papers 제목 키워드 기반

Over-complete representations on recurrent neural networks can support persistent percepts

2010-12-01 · NeurIPS 2010 12 · Shaul Druckmann, Dmitri B. Chklovskii

A striking aspect of cortical neural networks is the divergence of a relatively small number of input channels from the peripheral sensory apparatus into a large number of cortical neurons, an over-complete representatio…

PK-ICR: Persona-Knowledge Interactive Context Retrieval for Grounded Dialogue

2023-02-13 · Minsik Oh, Joosung Lee, Jiwei Li, Guoyin Wang

Identifying relevant persona or knowledge for conversational systems is critical to grounded dialogue response generation. However, each grounding has been mostly researched in isolation with more practical multi-context…

Data AugmentationResponse GenerationRetrieval

MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection

2026-04-30 · Weihai Lu, Zhejun Zhao, Yanshu Li, Huan He arxiv

Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting signals, remains challenging. Existing methods often face difficulties w…

Stance Detection

Compositional Image-Text Matching and Retrieval by Grounding Entities

2025-05-04 · Madhukar Reddy Vongala, Saurabh Srivastava, Jana Košecká

Vision-language pretraining on large datasets of images-text pairs is one of the main building blocks of current Vision-Language Models. While with additional training, these models excel in various downstream tasks, inc…

Image CaptioningImage-text matchingQuestion AnsweringRetrieval+3

Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs

2025-10-12 · Suyang Xi, Chenxi Yang, Hong Ding, Yiqing Ni 외 arxiv

Multimodal large language models (MLLMs) often fail in fine-grained visual question answering, producing hallucinations about object identities, positions, and relations because textual queries are not explicitly anchore…

Visual Question AnsweringMultimodal Reasoning