Tell Me What You See: Text-Guided Real-World Image Denoising
Image reconstruction from noisy sensor measurements is challenging and many methods have been proposed for it. Yet, most approaches focus on learning robust natural image priors while modeling the scene's noise statistics. In extremely low-light conditions, these methods often remain insufficient. Additional information is needed, such as multiple captures or, as suggested here, scene description. As an alternative, we propose using a text-based description of the scene as an additional prior, something the photographer can easily provide. Inspired by the remarkable success of text-guided diffusion models in image generation, we show that adding image caption information significantly improves image denoising and reconstruction for both synthetic and real-world images.
Code (0)
등록된 구현이 없습니다.
Tasks
DenoisingImage DenoisingImage GenerationImage ReconstructionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
Time series forecasting is critical to real-world decision making, yet most existing approaches remain unimodal and rely on extrapolating historical patterns. While recent progress in large language models (LLMs) highlig…
Time Series ForecastingDecision MakingDo You See What I See: Using Augmented Reality and Artificial Intelligence
Do You See What I See: Using Augmented Reality and Artificial Intelligence We explore the challenges of real-world applications of augmented reality (AR) and artificial intelligence (AI) through experiments that demons…
Keypoint EstimationPose EstimationSearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite. Recent work explores a pa…
What is an intelligent system?
The concept of intelligent system has emerged in information technology as a type of system derived from successful applications of artificial intelligence. The goal of this paper is to give a general description of an i…
V?: Guided Visual Search as a Core Mechanism in Multimodal LLMs
When we look around and perform complex tasks how we see and selectively process what we see is crucial. However the lack of this visual search mechanism in current multimodal LLMs (MLLMs) hinders their ability to fo…
Visual GroundingWorld Knowledge