DECDM: Document Enhancement using Cycle-Consistent Diffusion Models
The performance of optical character recognition (OCR) heavily relies on document image quality, which is crucial for automatic document processing and document intelligence. However, most existing document enhancement methods require supervised data pairs, which raises concerns about data separation and privacy protection, and makes it challenging to adapt these methods to new domain pairs. To address these issues, we propose DECDM, an end-to-end document-level image translation method inspired by recent advances in diffusion models. Our method overcomes the limitations of paired training by independently training the source (noisy input) and target (clean output) models, making it possible to apply domain-specific diffusion models to other pairs. DECDM trains on one dataset at a time, eliminating the need to scan both datasets concurrently, and effectively preserving data privacy from the source or target domain. We also introduce simple data augmentation strategies to improve character-glyph conservation during translation. We compare DECDM with state-of-the-art methods on multiple synthetic data and benchmark datasets, such as document denoising and {\color{black}shadow} removal, and demonstrate the superiority of performance quantitatively and qualitatively.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationDenoisingDocument EnhancementOptical Character RecognitionOptical Character Recognition (OCR)Shadow RemovalTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Unified Image Restoration and Enhancement: Degradation Calibrated Cycle Reconstruction Diffusion Model
Image restoration and enhancement are pivotal for numerous computer vision applications, yet unifying these tasks efficiently remains a significant challenge. Inspired by the iterative refinement capabilities of diffusio…
DenoisingImage DeblurringImage DehazingImage Inpainting+5Cycle-Consistent Speech Enhancement
Feature mapping using deep neural networks is an effective approach for single-channel speech enhancement. Noisy features are transformed to the enhanced ones through a mapping network and the mean square errors between …
Multi-Task LearningSpeech EnhancementA Two-stage Complex Network using Cycle-consistent Generative Adversarial Networks for Speech Enhancement
Cycle-consistent generative adversarial networks (CycleGAN) have shown their promising performance for speech enhancement (SE), while one intractable shortcoming of these CycleGAN-based SE systems is that the noise compo…
DenoisingSpeech EnhancementSpeech Enhancement Based on Cyclegan with Noise-informed Training
Cycle-consistent generative adversarial networks (CycleGAN) were successfully applied to speech enhancement (SE) tasks with unpaired noisy-clean training data. The CycleGAN SE system adopted two generators and two discri…
Speech EnhancementFrom Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning
Low-level enhancement and high-level visual understanding in low-light vision have traditionally been treated separately. Low-light enhancement improves image quality for downstream tasks, but existing methods rely on ph…
Zero-shot GeneralizationSemantic SegmentationDomain AdaptationImage Generation