Robust Learning of Diffusion Models with Extremely Noisy Conditions
Conditional diffusion models have the generative controllability by incorporating external conditions. However, their performance significantly degrades with noisy conditions, such as corrupted labels in the image generation or unreliable observations or states in the control policy generation. This paper introduces a robust learning framework to address extremely noisy conditions in conditional diffusion models. We empirically demonstrate that existing noise-robust methods fail when the noise level is high. To overcome this, we propose learning pseudo conditions as surrogates for clean conditions and refining pseudo ones progressively via the technique of temporal ensembling. Additionally, we develop a Reverse-time Diffusion Condition (RDC) technique, which diffuses pseudo conditions to reinforce the memorization effect and further facilitate the refinement of the pseudo conditions. Experimentally, our approach achieves state-of-the-art performance across a range of noise levels on both class-conditional image generation and visuomotor policy generation tasks.The code can be accessible via the project page https://robustdiffusionpolicy.github.io
Code (0)
등록된 구현이 없습니다.
Tasks
Conditional Image GenerationSimilar Papers 제목 키워드 기반
Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule
Recently, diffusion-based generative models have demonstrated remarkable performance in speech enhancement tasks. However, these methods still encounter challenges, including the lack of structural information and poor p…
Speech EnhancementDiffusion in the Dark: A Diffusion Model for Low-Light Text Recognition
Capturing images is a key part of automation for high-level tasks such as scene text recognition. Low-light conditions pose a challenge for high-level perception stacks, which are often optimized on well-lit, artifact-fr…
Image ReconstructionScene Text RecognitionBone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models
Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective inte…
Speech EnhancementTell Me What You See: Text-Guided Real-World Image Denoising
Image reconstruction from noisy sensor measurements is challenging and many methods have been proposed for it. Yet, most approaches focus on learning robust natural image priors while modeling the scene's noise statistic…
DenoisingImage DenoisingImage GenerationImage ReconstructionConsistent Diffusion Meets Tweedie: Training Exact Ambient Diffusion Models with Noisy Data
Ambient diffusion is a recently proposed framework for training diffusion models using corrupted data. Both Ambient Diffusion and alternative SURE-based approaches for learning diffusion models from corrupted data resort…
Memorization