paper-with-me

홈 › Papers

DONUT-hole: DONUT Sparsification by Harnessing Knowledge and Optimizing Learning Efficiency

2023-11-09 · Azhar Shaikh, Michael Cochez, Denis Diachkov, Michiel de Rijcke, Sahar Yousefi

This paper introduces DONUT-hole, a sparse OCR-free visual document understanding (VDU) model that addresses the limitations of its predecessor model, dubbed DONUT. The DONUT model, leveraging a transformer architecture, overcoming the challenges of separate optical character recognition (OCR) and visual semantic understanding (VSU) components. However, its deployment in production environments and edge devices is hindered by high memory and computational demands, particularly in large-scale request services. To overcome these challenges, we propose an optimization strategy based on knowledge distillation and model pruning. Our paradigm to produce DONUT-hole, reduces the model denisty by 54\% while preserving performance. We also achieve a global representational similarity index between DONUT and DONUT-hole based on centered kernel alignment (CKA) metric of 0.79. Moreover, we evaluate the effectiveness of DONUT-hole in the document image key information extraction (KIE) task, highlighting its potential for developing more efficient VDU systems for logistic companies.

📄 PDF Abstract BibTeX arXiv:2311.05778

Code (0)

등록된 구현이 없습니다.

Tasks

document understandingKey Information ExtractionKnowledge DistillationOptical Character RecognitionOptical Character Recognition (OCR)

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

OCR-Enhanced Multimodal ASR Can Read While Listening

2026-01-26 · Junli Chen, Changli Tang, Yixuan Li, Guangzhi Sun 외 arxiv

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve s…

Audio-Visual Speech RecognitionKnowledge Distillation

Donut Regression Discontinuity Designs

2023-08-28 · Cladia Noack, Chistoph Rothe

We study the econometric properties of so-called donut regression discontinuity (RD) designs, a robustness exercise which involves repeating estimation and inference without the data points in some area around the treatm…

regression

Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document

2025-09-30 · Adnan Ben Mansour, Ayoub Karine, David Naccache arxiv

Recent advances in Visually-rich Document Understanding rely on large Vision-Language Models like Donut, which perform document-level Visual Question Answering without Optical Character Recognition. Despite their effecti…

Visual Question AnsweringKnowledge DistillationModel Compression

DONUT: Physics-aware Machine Learning for Real-time X-ray Nanodiffraction Analysis

2025-07-18 · Aileen Luo, Tao Zhou, Ming Du, Martin V. Holt 외 arxiv

Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis …

OCR-free Document Understanding Transformer

2021-11-30 · Geewook Kim, Teakgyu Hong, Moonbin Yim, Jeongyeon Nam 외

Understanding document images (e.g., invoices) is a core but challenging task since it requires complex functions such as reading text and a holistic understanding of the document. Current Visual Document Understanding (…

Document Image Classificationdocument understandingKey-value Pair ExtractionOptical Character Recognition+2