paper-with-me

홈 › Papers

UDoc-GAN: Unpaired Document Illumination Correction with Background Light Prior

2022-10-15 · Yonghui Wang, Wengang Zhou, Zhenbo Lu, Houqiang Li

Document images captured by mobile devices are usually degraded by uncontrollable illumination, which hampers the clarity of document content. Recently, a series of research efforts have been devoted to correcting the uneven document illumination. However, existing methods rarely consider the use of ambient light information, and usually rely on paired samples including degraded and the corrected ground-truth images which are not always accessible. To this end, we propose UDoc-GAN, the first framework to address the problem of document illumination correction under the unpaired setting. Specifically, we first predict the ambient light features of the document. Then, according to the characteristics of different level of ambient lights, we re-formulate the cycle consistency constraint to learn the underlying relationship between normal and abnormal illumination domains. To prove the effectiveness of our approach, we conduct extensive experiments on DocProj dataset under the unpaired setting. Compared with the state-of-the-art approaches, our method demonstrates promising performance in terms of character error rate (CER) and edit distance (ED), together with better qualitative results for textual detail preservation. The source code is now publicly available at https://github.com/harrytea/UDoc-GAN.

📄 PDF Abstract BibTeX arXiv:2210.08216

Code (1)

harrytea/udoc-gan 공식 구현 pytorch

Similar Papers 제목 키워드 기반

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System

2025-02-14 · Karan Taneja, Ashok K. Goel

Multimodal AI is an important step towards building effective tools to leverage multiple modalities in human-AI communication. Building a multimodal document-grounded AI system to interact with long documents remains a c…

AI AgentResponse Generation

DocTr: Document Image Transformer for Geometric Unwarping and Illumination Correction

2021-10-25 · Hao Feng, Yuechen Wang, Wengang Zhou, Jiajun Deng 외

In this work, we propose a new framework, called Document Image Transformer (DocTr), to address the issue of geometry and illumination distortion of the document images. Specifically, DocTr consists of a geometric unwarp…

Optical Character Recognition (OCR)

Synthetic Document Question Answering in Hungarian

2025-05-29 · Jonathan Li, Zoltan Csaki, Nidhi Hiremath, Etash Guha 외

Modern VLMs have achieved near-saturation accuracy in English document visual question-answering (VQA). However, this task remains challenging in lower resource languages due to a dearth of suitable training and evaluati…

Optical Character Recognition (OCR)Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Logic Error Localization in Student Programming Assignments Using Pseudocode and Graph Neural Networks

2024-10-11 · Zhenyu Xu, Kun Zhang, Victor S. Sheng

Pseudocode is extensively used in introductory programming courses to instruct computer science students in algorithm design, utilizing natural language to define algorithmic behaviors. This learning approach enables stu…

DiagnosticGraph Neural Network

Document Rectification and Illumination Correction using a Patch-based CNN

2019-09-20 · Xiaoyu Li, Bo Zhang, Jing Liao, Pedro V. Sander

We propose a novel learning method to rectify document images with various distortion types from a single input image. As opposed to previous learning-based methods, our approach seeks to first learn the distortion flow …

Optical Character Recognition (OCR)