paper-with-me

홈 › Papers

Robust Learning of Diffusion Models with Extremely Noisy Conditions

2025-10-11 · Xin Chen, Gillian Dobbie, Xinyu Wang, Feng Liu, Di Wang, Jingfeng Zhang arxiv

Conditional diffusion models have the generative controllability by incorporating external conditions. However, their performance significantly degrades with noisy conditions, such as corrupted labels in the image generation or unreliable observations or states in the control policy generation. This paper introduces a robust learning framework to address extremely noisy conditions in conditional diffusion models. We empirically demonstrate that existing noise-robust methods fail when the noise level is high. To overcome this, we propose learning pseudo conditions as surrogates for clean conditions and refining pseudo ones progressively via the technique of temporal ensembling. Additionally, we develop a Reverse-time Diffusion Condition (RDC) technique, which diffuses pseudo conditions to reinforce the memorization effect and further facilitate the refinement of the pseudo conditions. Experimentally, our approach achieves state-of-the-art performance across a range of noise levels on both class-conditional image generation and visuomotor policy generation tasks.The code can be accessible via the project page https://robustdiffusionpolicy.github.io

📄 PDF Abstract BibTeX arXiv:2510.10149

Code (0)

등록된 구현이 없습니다.

Tasks

Conditional Image Generation

Similar Papers 제목 키워드 기반

Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule

2024-09-08 · Siyi Wang, Siyi Liu, Andrew Harper, Paul Kendrick 외

Recently, diffusion-based generative models have demonstrated remarkable performance in speech enhancement tasks. However, these methods still encounter challenges, including the lack of structural information and poor p…

Speech Enhancement

Diffusion in the Dark: A Diffusion Model for Low-Light Text Recognition

2023-03-07 · Cindy M. Nguyen, Eric R. Chan, Alexander W. Bergman, Gordon Wetzstein

Capturing images is a key part of automation for high-level tasks such as scene text recognition. Low-light conditions pose a challenge for high-level perception stacks, which are often optimized on well-lit, artifact-fr…

Image ReconstructionScene Text Recognition

Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models

2026-01-18 · Sina Khanagha, Bunlong Lay, Timo Gerkmann arxiv

Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective inte…

Speech Enhancement

Tell Me What You See: Text-Guided Real-World Image Denoising

2023-12-15 · Erez Yosef, Raja Giryes

Image reconstruction from noisy sensor measurements is challenging and many methods have been proposed for it. Yet, most approaches focus on learning robust natural image priors while modeling the scene's noise statistic…

DenoisingImage DenoisingImage GenerationImage Reconstruction

Consistent Diffusion Meets Tweedie: Training Exact Ambient Diffusion Models with Noisy Data

2024-03-20 · Giannis Daras, Alexandros G. Dimakis, Constantinos Daskalakis

Ambient diffusion is a recently proposed framework for training diffusion models using corrupted data. Both Ambient Diffusion and alternative SURE-based approaches for learning diffusion models from corrupted data resort…

Memorization